AERIOXFLUX
← Science
Science · biotech

Claude Found a New Phage Enzyme System, Then Couldn't Find It Again

Anthropic's agents surfaced a CRISPR-like reverse transcriptase family in 21.5 hours without human help, and the preprint is unusually candid that ten reruns of the same campaign missed it.

Flux Desk·2026-09-24·5 min read

Most discoveries in genome mining start the same way. Someone looks at a stretch of sequence, notices something that should not be there, and decides it is worth a week of their life. Pipelines cannot do that part. They find what they were told to find.

On September 23, Anthropic said its Claude agents did that part on their own. A blog post and a preprint from the company's life sciences group describe a previously unreported family of enzymes in bacteriophages, the viruses that infect bacteria. Anthropic calls it array-associated reverse transcriptases, or ART. The preprint, "Autonomous AI agents discover reverse transcriptases with tandem repeat arrays," has not been peer reviewed.

What the agents actually did

Anthropic's researchers wrote a research brief and let a multi-agent harness run it. The brief asked for new reverse transcriptase (RT) systems, identified by new partner genes sitting next to the RT. RTs copy RNA into DNA, and bacterial versions such as retrons already underpin work in genome engineering.

According to the preprint, the agents ran on Claude Mythos 5. They built their own search profiles, pulled about 200,000 RT clusters out of 1.9 billion metagenomic protein clusters, sorted them into nine classes, sampled roughly 11,000 RT loci and scored 3,564 protein families that kept turning up nearby. Sixteen families passed the agents' own selection criteria. A seventeenth was promoted from one of 98 follow-up tasks the agents opened themselves.

The full campaign was 119 tasks and 949 agent sessions: 77 agent-hours and 215.6 million tokens over 21.5 hours of wall-clock time, which the authors say ran "without human intervention." It ended with 19 reports, ranked in a tournament-style evaluation. Anthropic's blog rounds these to about 950 agents, 210 million tokens, 21 hours and 20 candidates. The preprint's figures are the precise ones.

The result was not a flood of breakthroughs. Of the 17 candidate partner families, only three held up as previously unreported RT associations. The other 14 were annotation artifacts, parts of known systems, or genes that simply lived nearby. Three flagged RTs turned out to be new lineages. That hit rate reads like real screening work, not a press release.

The anomaly nobody asked for

ART was not in the brief. The brief asked about partner proteins. One worker, reading raw DNA upstream of an odd RT in a jumbo phage, noticed a run of evenly spaced repeats and flagged it. The resulting system has three parts: the RT, a dedicated partner gene, and a long array of roughly 200-nucleotide non-coding repeat units. The layout resembles a CRISPR array, which is why the comparisons started immediately.

Anthropic is clear on what is not new. The enzyme itself had been seen before in jumbo phages. The preprint gives one example: the original genome report for the MarsHill phage identified the RT but described neither the repeats nor the partner gene. Claude appears to be the first to connect the three.

Follow-up work was split between machine and human. A Claude Science session found and reanalyzed a published RNA-sequencing study of Staphylococcus phage SA1, which carries an ART system. At 15 minutes after infection, RNA from the array made up as much as 8% of phage RNA, among the most abundant transcripts. Anthropic's human scientists then expressed the SA1 system on plasmids in E. coli and saw the array break into discrete short RNAs. The company says all lab work is done by people, with Claude helping interpret the data.

The rerun problem

The most useful paragraph in the preprint is the one that works against the headline. The authors ran the same campaign ten more times. Nearly every run sampled ART loci, and in two runs workers followed up on the lineage. None of them read the upstream DNA, and the array was missed every time. The authors blame the broad search space and the harness's non-deterministic behavior.

So they built fixed benchmarks instead. Seven Claude models were asked to describe the ART system, scored by a judge model against ten features. Four models (Opus 5.5, Mythos 5.1, Mythos 5 and Opus 5) clearly outperformed three older ones (Opus 4.6, Opus 4.8 and Sonnet 5). With the DNA pasted straight into context, the strongest models described the array in at least 90% of attempts. Given the same data as files with tools, performance fell as low as 32% for Opus 5. In 39% of those attempts, the model never read a continuous 200-nucleotide stretch, so it never saw more than about one repeat.

That is a practical lesson for anyone building research agents. More tools made the models worse at the one observation that mattered, because the tools let them skip reading the data. The discovery depended on an agent actually looking.

The preprint goes one step further. Using Anthropic's interpretability methods on the original session transcript, the authors report internal signals in Mythos 5 that respond to repeated DNA, and argue this "genomic vision" enabled the find. That is a claim about Anthropic's own model, tested by Anthropic, and worth treating as such.

What is known and what is not

Feng Zhang of MIT and the Broad Institute, one of the scientists behind CRISPR gene editing, reviewed an early copy. The Next Web quoted him calling the RNA-repeat arrays "genuinely intriguing." Independent labs have not yet reproduced anything.

The preprint is direct about the gaps. The authors have not shown that the RT is active, that the array RNAs are its substrates, whether the RT and partner interact, or what the system does for the phage. Unlike CRISPR spacers, the ART units are conserved between related phages, and they are transcribed far in excess of the enzyme. The authors read that as one enzyme supplied with a repertoire of templates or baits. That is a hypothesis, not a function.

It also lands days after three papers in Science described VIPR, another phage-encoded system with CRISPR-like traits. The field is plainly rich, and the question is who finds the next one first.

Why it matters

The claim here is narrower than "AI discovers biology" and more interesting. The agents supplied both the expectation and the judgment that something broke it, which is the step that usually needs a senior scientist. They also missed it ten times out of eleven, and handing the same data over as files instead of in context sharply cut how often models spotted the array.

For labs, the takeaway is concrete. A 21-hour, 215-million-token sweep of a database this size is short enough to run more than once, and the rerun data says you should. Run it several times, make the agents read the raw sequence, and keep the bench as the final judge.

#anthropic#bacteriophage#reverse-transcriptase#genome-mining#ai-science

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.