memujo
AI8 min read

Claude Found an Enzyme System, Then Missed It 10x

Anthropic's agent swarm burned 215 million tokens finding an unknown enzyme system, then failed to reproduce it ten times. The funnel tells you why.

By Alice

In this article
  1. 01What Was Actually Found
  2. 02The Search Funnel, Number by Number
  3. 03Our Read: This Is a Retrieval Problem Wearing a Biology Lab Coat
  4. 04Outlook

On September 23, Anthropic announced that Claude agents, given a single prompt to hunt for new reverse transcriptases in a DNA sequence database, autonomously flagged a previously uncharacterized enzyme system in bacteriophage DNA. The system sits next to a long array of evenly spaced DNA repeats, a layout that recalls CRISPR, and Anthropic says it does not yet know what it does. That alone would be the story. The more interesting number arrived in the preprint: Anthropic ran the same campaign ten more times, and all ten reruns missed the discovery entirely.

A finding that a model produced once and could not reproduce ten times is not a contradiction, and it is not a debunking. It is a measurement of something engineers care about deeply and science communicators usually flatten: the gap between capability and reliability. Read carefully, the announcement is as much a report on agent search behavior as it is a biology result.

What Was Actually Found

The new life sciences group at Anthropic, formed in spring 2026 according to the company's announcement, gave Claude one research brief: search a large protein cluster database for interesting new reverse transcriptases, the enzymes that copy RNA into DNA. Scientists were involved at two points only, the initial prompt and the lab work afterward. In between, Claude agents explored the database, investigated RT families, and used their own judgment about which candidates were worth a report.

The system the agents surfaced is what the team calls array-associated reverse transcriptases, or ART. It has three components: a reverse transcriptase, a partner gene of unknown function beside it, and an array of evenly spaced DNA repeats. The repeat array is the CRISPR echo. In CRISPR systems, a comparable array stores a bank of RNA guides that make the system programmable, which is exactly why CRISPR became a gene editing tool rather than a curiosity.

Three honest caveats come straight from the sources, and Anthropic states all of them. First, the enzyme itself was not new. Earlier studies had identified the reverse transcriptase in jumbo phages; one of those genome reports describes the enzyme without mentioning the repeats or the partner gene. Claude appears to be the first to notice the association. Second, the function is unknown. Lab tests showed the array is expressed as a set of distinct short RNAs, and in published data from one phage those RNAs reached as much as 8 percent of total phage RNA 15 minutes after infection, per the preprint as reported by The Next Web. But the team has not yet shown the enzyme is active, or that it acts on those RNAs. Third, the preprint has not been peer reviewed.

External reaction has been warm but bounded. Feng Zhang, the CRISPR pioneer at MIT and the Broad Institute, reviewed the preprint and called the identification of RNA repeat arrays associated with reverse transcriptases "genuinely intriguing" and worthy of further investigation. Stanley Qi, a bioengineering professor at Stanford, told Al Jazeera that what stands out is the ability to recognize an unusual biological pattern that was difficult to detect before. Note the shape of both endorsements: they validate the pattern recognition, not a therapeutic claim. Reuters reported earlier this month that Anthropic had quietly stood up the biology lab behind this work, so the announcement is a result from a deliberate in-house program, not a one-off demo.

The Search Funnel, Number by Number

Anthropic and the preprint release a full set of funnel numbers, which is more transparency than most AI-for-science announcements offer. Assembling them into one table is the most useful thing a reader can do with this story, because the funnel, not the headline, shows what kind of method this is.

Stage Count Notes
Protein clusters searched 1.9 billion Database per the preprint
Agent sessions 949 Anthropic's blog rounds to "roughly 950"
Wall-clock runtime 21.5 hours Unattended, per the preprint
Tokens consumed 215.6 million Blog rounds to 210 million
RT clusters gathered ~200,000 Recovered by the agents
Candidate partner families scored 3,564 Blog rounds to 3,500
Reports filed for human review 19 Blog says "the 20 most-compelling"
Lab-validated new systems 1 ART
Successful reruns of the campaign 0 of 10 Same brief, same models

Two derived numbers are worth sitting with. The campaign averaged roughly 10 million tokens per hour and about 227,000 tokens per agent session, so this was not one long context window reading the genome of the world; it was many short-lived, tool-using sessions fanning out. And the funnel narrows by about four orders of magnitude per stage, from 1.9 billion clusters to one reportable system, which is what rare-event search looks like when you write it down.

Our Read: This Is a Retrieval Problem Wearing a Biology Lab Coat

Strip away the pipettes and this campaign is an information retrieval system with a stochastic searcher at the top of the stack. A corpus of 1.9 billion items, a recall target nobody can enumerate in advance ("interesting"), a ranking function implemented as an LLM's judgment, and a human labeler (the scientist) who only ever sees the top of the list. Every design choice follows from that framing, and so does every weakness.

The rerun result is the most valuable data point in the preprint precisely because it separates two failure modes that get conflated in coverage. Anthropic's own diagnostic, as reported by The Next Web, is the clean one: when the four most capable models were handed the relevant DNA directly, they described the array correctly in at least 90 percent of attempts. When the same models had to work through files and tools, the success rate fell as low as 32 percent, often because the agent never read enough of the sequence to see a full repeat. The model could do the task. The agent scaffold could not reliably put the right bytes in front of the model. That is a context management bug with a biology costume, and most "the model isn't capable enough" conclusions in agent systems are actually this same bug.

The 0-for-10 reruns then follow from search geometry rather than incompetence. Finding ART required an agent to wander into the upstream region of one particular gene, an act the original run performed as a side path, prompted by the agent noticing something "spectacular" while reading raw DNA beside an unusual enzyme. A free-roaming agent's trajectory through 1.9 billion clusters is a sample from a path distribution. The discovery event sits in a thin tail of that distribution. Ten draws missing a thin-tail event is not evidence the event is impossible; it is an upper bound on its per-run probability somewhere below roughly 10 percent, and an indictment of an exploration strategy with no coverage guarantee. A deterministic k-mer or repeat-detection scan over the same corpus, the classic bioinformatics approach, would find repeat arrays by construction because it has to look everywhere. The agents won because judgment beat brute force on the ranking stages, and lost on recall because nobody forced the searcher to be exhaustive.

The economics angle is surprisingly favorable despite that. Call it 215.6 million tokens. Even at a blended frontier-model price of, say, $10 per million tokens, a number I am assuming rather than quoting, since Anthropic has not disclosed rates, the compute for the entire campaign lands near $2,200. Anthropic's own comparison is that this class of analysis takes an expert scientist weeks to months. Whether or not you trust the price estimate, the cost asymmetry is the story: a few thousand dollars of tokens against months of a skilled scientist's time is a ratio that justifies running the campaign ten more times just to catch what the first run missed. Compare that to the hardware-scale bets behind models, like the 2,304 GPU loan this site covered last week, and agent-discovery campaigns look like rounding error in spend with a genuinely open-ended payoff.

There is a second-order point Anthropic volunteers, and it may matter more than ART. With hundreds to thousands of candidate reports per campaign, the hypotheses themselves have become an object of study: the team now researches what distinguishes proposals worth testing from those they set aside, and feeds that back as instructions to teach Claude something like scientific taste. Any engineer who has built a recommendation or alerting system recognizes this loop. Once generation is cheap, triage is the product, and the scarce human hours move entirely to the ranking boundary. That is also where this whole category of AI science lives or dies, because a pipeline that produces 3,564 candidates and one winner is only as good as the scientist's willingness to keep reading reports.

We covered the general problem of trusting vendor-reported numbers in Reading AI Benchmark Claims Like a Skeptic, and this announcement is a case study in doing it right: every number above is Anthropic's own, the limits are stated by Anthropic, and the one external voice (Zhang) comments on evidence, not implications. The Mythos 5 release that supplied the models for this campaign is the other half of the story: frontier labs are now differentiating on what their agents can be pointed at, and biology is the highest-status target available.

Outlook

What would change this assessment, concretely. If Anthropic or an independent group publishes a demonstration that the ART enzyme is catalytically active, and especially if the repeat array is shown to guide that activity the way CRISPR guides do, this jumps from intriguing pattern to a candidate gene editing mechanism, which is the framing Dario Amodei used on X and that should be treated as a hypothesis, not a finding. If instead years of work show the array is a phage artifact with no enzymatic role, the retrieval lesson survives anyway, because the scaffold failure data (90 percent versus 32 percent) is useful regardless of what ART turns out to be.

For builders, the engineering takeaways are available now. First, evaluate agent systems at two layers separately, model-with-context and model-with-tools, because the delta between them is where your reliability actually lives. Second, if your agent searches a large space for rare events, give it coverage guarantees, scheduled sweeps, forced upstream-region reads, things that a random walk will not provide at ten reruns. Third, publish your rerun failures. Anthropic losing 10 of 10 reruns and printing that number does more for the credibility of agentic discovery than any success montage, and it hands every other team the one metric that matters: not can the agent find it, but how often.

  • #anthropic
  • #ai-agents
  • #ai4science
  • #benchmarks
  • #reproducibility
  • #genomics

Sources

Share this story