When the Machine Notices First
Claude’s ART discovery changes the Frontier question: what happens when an AI notices something humans missed, human judgement decides it matters, and physical reality determines what survives?
When I started THE FRONTIER, the picture in my head was still mostly human-first. A person notices something, asks a question or brings years of experience into an interaction with increasingly capable AI, and the machine takes that insight somewhere the person could not realistically take it alone. That was already enough to make me ask which part of the interaction produced the useful intelligence, but recent research from Anthropic has complicated the picture because, according to the company's account, this time the human did not appear to notice the important thing first. Claude did.
In a recent preprint, Anthropic describes giving Claude a broad research brief: search genomic databases for previously uncharacterised systems associated with reverse transcriptases, enzymes that copy RNA into DNA. The human researchers established that direction and would later carry out the laboratory work, but the computational exploration proceeded without human intervention. Over 21.5 hours, roughly 950 Claude agent sessions screened nearly 200,000 reverse-transcriptase clusters, investigated thousands of possible partner families and produced nineteen reports for comparison. During that search, one agent encountered an unusual reverse transcriptase, examined the DNA around it and found a long repeating array alongside a neighbouring partner gene. It counted the repeats, compared the arrangement with known systems, searched the literature and pursued the anomaly far enough for the result to reach human review. Anthropic now calls the family array-associated reverse transcriptases, or ART.
The caution belongs near the beginning because we still do not know what ART does. The underlying reverse transcriptase had already appeared in earlier research; Anthropic's reported novelty concerns the proposed system combining that enzyme, the repeat array and the partner protein. The work has not yet been peer reviewed, and the resemblance to CRISPR is architectural rather than functional: both contain repeated DNA associated with a potentially programmable molecular system, but ART has not been shown to edit genes or perform any comparable operation. Public infection data and Anthropic's laboratory expression experiments indicate that ART arrays are transcribed and can appear as distinct short RNAs, yet the system's biological function remains unresolved. Further experiments may show that Claude found something important, something narrow or simply something unusual.
Even with that uncertainty, something relevant to THE FRONTIER has already happened. The researchers did not point Claude towards the repeat array and ask it to explain what the pattern meant. They gave the system a much broader search problem, and one agent pursued a feature outside the protein-partner search that had originally organised the campaign. Look closely at that moment: the agent read raw DNA beside the reverse transcriptase, recognised an atypical repetition and extended the investigation. Calling this human-style curiosity would add more than the evidence allows, but describing it only as a database search removes the part that makes the case interesting. The search did not begin with the correct question already specified. A candidate feature became a question because the system selected it for further attention, which brings us to the distinction I want this paper to investigate: is noticing meaningfully different from searching?
Scientific discovery often begins before anyone knows the correct question. A strange pattern appears, something refuses to behave as expected or two things that were never considered together begin to look related. The scientist does not start by knowing what the discovery is; the scientist decides that an anomaly deserves another look. Anthropic's account suggests that Claude performed a functional version of that selection inside this campaign, but it also shows how much human judgement shaped the conditions in which the selection became possible. Claude can produce hundreds or thousands of candidate reports, so the researchers study which ones they judge worth testing and feed those preferences back into later instructions, teaching the system to imitate what Anthropic calls their own scientific taste. When a candidate survives that process, human scientists decide whether it deserves scarce laboratory time, perform the experiments and interpret what comes back with Claude's help.
The resulting loop no longer fits neatly into the sequence I first imagined: human insight followed by machine-scale intelligence. Human judgement shapes the machine's search, the machine extends that judgement across a space the humans cannot inspect themselves, an agent identifies something the humans had not identified, the finding returns to human judgement, and physical experiments provide evidence that changes what can be investigated next. In the ART case, the sequence looks more like machine observation, human judgement, physical experiment and further interpretation, while each stage depends on work done at the others. The process carries a developing form of scientific judgement across its participants, although I am not yet ready to say whether that judgement belongs to the interaction itself or whether the interaction is simply moving human and machine capabilities into the places where each can be useful.
PACMAN, a framework developed by researchers at Princeton Plasma Physics Laboratory and Princeton University, gives us another version of that structure. The researchers tested it in five experiments on the DIII-D fusion facility, with humans setting the objectives and hardware safety limits while machine-learning models responded to plasma measurements in milliseconds. In one experiment, a model predicted a damaging tearing-mode instability about 200 milliseconds before it formed, allowing the system to change the plasma first. This is not the same kind of noticing as the ART search, but it strengthens the argument that physical reality may participate in the loop rather than merely judge an answer at the end. The plasma continuously produces new measurements, those measurements change the machine's next action, and human researchers later review what happened and adjust the controllers. Reality is supplying information throughout the process.
Paper2Agent expands the possible interacting unit in another direction. Published in Nature, the framework converts papers and their associated code, data and methods into agents that can answer questions, reproduce analyses and apply methods to new problems. In one case study, agents derived from three different bodies of genomics research worked together to prioritise GPR137 as a probable causal gene at a psoriasis-associated locus, while a human researcher selected the validation strategy from those the system proposed. The authors are careful to keep open-ended hypothesis generation and evaluation human-in-the-loop, but the case still matters for this investigation because knowledge produced by different groups of scientists can become active computational participants in the same analysis. The relevant interaction may therefore extend beyond one human and one model to include several agents, several bodies of human knowledge, datasets, tools and the judgement that determines which direction deserves pursuit.
By this point, asking which AI discovered this? or even which scientist discovered this? may be too narrow for some of the systems we are constructing. A scientist may define the objective, other scientists may have produced the knowledge encoded in several papers, different agents may apply that knowledge, an instrument or database may supply evidence, one agent may flag an anomaly, another may challenge the interpretation, a laboratory may reject the hypothesis and a human may recognise something useful in the failure. The parts remain identifiable and their contributions should not be blurred together, yet the result can become difficult to explain by examining any one of them alone. That possibility is worth investigating, but it is not enough to declare that a new collective intelligence has appeared.
The AIntibody Challenge provides useful resistance to that temptation. In a prospective, blinded benchmark published in Nature Biotechnology, researchers experimentally tested 511 AI-designed or AI-ranked antibodies from twenty-nine organisations across three tasks. Several groups produced developable antibodies with very high affinity, particularly in a defined affinity-maturation task, but those successes did not transfer consistently. In the task that asked models to rank high-affinity antibodies within sequence clusters, every model except one performed worse than random selection, while many out-of-library designs failed to outperform standard experimental selection. The study examined a single, unusually well-characterised antigen under favourable data conditions, and its authors describe the results as an upper bound for that target class rather than evidence of general scientific capability.
The only reason those successes and failures can be separated so clearly is that the computational outputs were forced through standardised physical experiments. That matters because the frontier we are investigating cannot be defined as the place where an AI produces something impressive. These cases are more useful when they reveal how the participants succeed and fail differently. Claude can inspect a vast genomic search space and identify a pattern that human researchers had not characterised as a complete system, yet humans still decide whether the pattern deserves experimental resources. A model can design an antibody that looks exceptional computationally and still fail when physical biology tests it. A human scientist can bring decades of experience and still miss a relationship distributed across hundreds of thousands of sequences. Reality can answer one question and create three more. The possibility I want to keep open is not that every interaction becomes intelligent, but that different insufficiencies may sometimes be connected in a way that allows the larger process to reach something none of its parts could reach efficiently alone.
This brings me back to ART, because the strongest question raised by the case is not simply whether Claude discovered something. Anthropic's account supports the narrower claim that the computational lead emerged from Claude's exploration, while the human scientists supplied the research direction, chose what merited laboratory work, performed the experiments and remain responsible for establishing biological significance. The harder question is what made that lead possible. Was it the capability of the model, the scale of roughly 950 agent sessions working in parallel, the scientists' judgement encoded upstream, the biological knowledge in the literature, or the particular path one agent followed when it looked beyond the reverse transcriptase and examined the surrounding DNA? The useful explanation may require all of them, but we should not assume that before we know how much each contributed.
The preprint adds a detail that makes the question sharper. Anthropic ran the same campaign ten more times. Nearly every completed run sampled ART loci, and two launched follow-up investigations into the lineage, yet none read the relevant upstream DNA and every rerun missed the array. The authors attribute this to the breadth of the search space and the non-deterministic behaviour of the agent harness. That result does not prove emergence, and it may tell us as much about the fragility of the workflow as it does about scientific judgement, but it suggests that the frontier may depend not only on how much intelligence is available to a system. It may also depend on where the interaction looks and which path it takes through the possibilities in front of it.
Scientists with similar knowledge can notice different things, and two researchers can observe the same anomaly while only one decides not to dismiss it. Discovery has always depended partly on recognising that something ordinary-looking may not be ordinary at all. If agentic systems begin performing a reliable functional version of that behaviour—not merely answering a question but selecting which unexplained feature deserves a new one—then AI's role in discovery has changed. Yet the failed reruns also warn us not to confuse one productive trajectory with a stable capability. Before calling this machine curiosity or scientific intuition, we need to know whether the behaviour can be reproduced, what cues caused the system to pursue the array and how much of that apparent judgement came from human scientific taste encoded before the search began.
I am therefore not ready to give the noticing to Claude alone. Outside scientific reaction has also raised questions about how much of the discovery should be attributed to Claude, including whether researchers had previously studied the same system, while its biological function remains unresolved. It operated inside an environment designed by humans, used knowledge produced by humans, followed a direction chosen by humans and returned its findings to humans who could decide whether the laboratory should care. At the same time, I do not think the AI searched a database is an adequate account of what happened. Inside a vast search space, the agent identified an anomaly that had escaped characterisation as a complete system and pursued it far enough to change what the human scientists considered worth testing. Whether that was elicitation of a capability already available within the model, an unusually successful search path or something that emerged from the trajectory of the wider process remains unresolved.
The question that opened THE FRONTIER was where exactly the useful intelligence appears, and ART now adds another: who decides what is worth noticing? If the human can supply the research direction and the machine can supply the first observation, if machine exploration changes human understanding and human judgement changes what the machine investigates, while physical reality continues correcting both, then the correct unit of analysis may sometimes be larger than either participant. Or we may eventually find that separating the contributions more carefully explains the result without requiring any larger intelligence at all.
ART has not resolved that argument, but it has changed the research. We now have a case in which the machine appears to have noticed first, a laboratory result that establishes less than the headline might suggest, ten failed reruns that make the path through the search space difficult to ignore, and a biological function we still do not know. The evidence therefore gives us a firmer direction without giving us a final answer: the frontier may depend not only on how much intelligence is available to a system, but on what gets noticed, where the interaction looks, which path it takes through the possibilities in front of it and what happens when human judgement and physical reality enter after the observation has been made. If the first Frontier paper asked whether the frontier may exist in the interaction, ART gives us a sharper place to look: the moment the machine notices first, the human decides it matters, and reality determines what survives.
— Teff