Cover: AI-generated editorial composition by TMRW. Findings are from Anthropic's announcement and its preprint, which has not been peer reviewed.
On 23 September, Anthropic said a swarm of Claude agents had found an enzyme system in viral DNA that nobody had described before. It has a feature associated with CRISPR, the gene-editing tool that won a Nobel Prize. What it does, nobody knows yet. That gap is the most honest part of the announcement, and the best guide to what AI is actually adding to biology right now.
What Claude found
The system lives mainly in bacteriophages, the viruses that infect bacteria. Anthropic calls it array-associated reverse transcriptases, or ART. It has three parts: a reverse transcriptase (an enzyme that copies RNA into DNA), a partner gene next to it, and a long run of evenly spaced repeated DNA.
That last part is why people are paying attention. In CRISPR systems, a repeat array holds a bank of different RNA sequences, and those RNAs are what make CRISPR programmable: you choose where it acts. Anthropic's first lab experiments show ART arrays are also read out as distinct short RNAs during infection of a Staphylococcus phage. That hints at something similar. It does not show it.
Anthropic is careful about what is new. The enzyme itself had been identified in earlier studies of a jumbo phage. The first genome report on one of these phages described the enzyme and missed the repeats and the partner gene. Claude appears to be the first to notice that the three belong together. The preprint also notes differences from CRISPR: no Cas genes nearby, and spacers of 120 to 220 nucleotides, far longer than CRISPR's.
How the search ran
Anthropic's scientists wrote one prompt: search a massive sequence database for interesting new reverse transcriptases. Roughly 950 Claude agents then worked for about 21 hours and used 210 million tokens. They gathered more than 200,000 reverse transcriptases across 1.9 billion protein clusters, scored 3,564 candidate families, and wrote human-readable reports on the 20 most compelling. The company says that kind of survey takes an expert weeks to months.
The find came from a detour. One agent was checking a side question and pulled the raw DNA beside an odd-looking enzyme into its context. It wrote: "[The DNA next to the RT] is spectacular: I can see by eye a tandem repeat array … that's a CRISPR-like … repeat array?!" Then it counted the repeats, measured their spacing, checked the literature and filed a report.
Two details in the preprint deserve more attention than the exclamation. First, the agents threw out most of their own ideas: of 17 candidate partner families, only three were confirmed as previously unreported, and the other 14 were rejected as annotation artifacts or parts of known systems. Second, Anthropic's benchmarks suggest the discovery depended on the model reading raw DNA rather than working only from annotations. A pipeline that summarizes sequences before a model sees them would probably have missed this.
What this does and does not prove
It proves that AI agents can now do the noticing step of genome mining at a scale no lab staff can match. In this field, that step is a real bottleneck. CRISPR itself was first spotted as an odd repeat in bacterial DNA, years before anyone knew what it did.
It does not prove that Claude discovered a useful tool. Plenty of odd systems in nature turn out to be dead ends. Feng Zhang of MIT and the Broad Institute, one of CRISPR's pioneers, reviewed the preprint and chose his words carefully: the link between reverse transcriptases and RNA-repeat arrays is "genuinely intriguing and merits further investigation." That is a scientist saying "worth a look," not "breakthrough."
Compare it with this month's AI math claims, which we covered in OpenAI's Millennium Prize announcement. A proof can be checked on paper. A protein has to be expressed, purified and tested, and that work is slow. At Anthropic, human scientists still do all of it, in a lab limited to biosafety levels 1 and 2.
Where the bottleneck moves
If agents can turn a database into 20 good leads in a day, the scarce resource becomes bench time and judgment about which leads deserve it. Anthropic says as much: it now studies Claude's hypotheses to learn which ones its scientists choose to test, and feeds that back into the instructions. That is the real product here, a triage system for experiments.
For a biotech lab sitting on unexplored sequence data, the lesson is concrete. Let the model read the raw sequence, make it try to refute its own candidates, and budget for the lab work that follows. The ART result will be judged in months, when someone learns what the system actually does. Until then, it is a strong lead, not a discovery with a use.
