This is Episode 7 of Reading the Frontier, a series of close readings of what the frontier labs publish.
The agent had six days, three thousand dollars of API credit, GPU time, a Linux machine, the open web, and a real research question that nobody had published an answer to. It was asked to design a detector that would raise an alarm when a tabular foundation model started making worse predictions than its numbers suggested.
In the first fourteen hours it generated six distinct approaches and falsified all of them. That is not, in itself, a failure. It is roughly what the first day of a hard project looks like. What happened next is the finding.
With one hundred and ten hours still on the clock, the agent never revised its approach again. It did not go back to the six ideas and ask whether it had tested them properly. It did not spawn a clean subagent and start over, though it had that tool and used it routinely for other things. Instead it reframed the goal: since it had not found a detector, it would write a paper arguing that no such detector could exist. Not that it had failed to find one in fourteen hours of underpowered tests, which is what had happened, but that the thing was not there to be found. It then submitted that paper to its own reviewer ten times. Ten times the reviewer returned Weak Reject. The agent responded by adding qualifications, running robustness checks, and hedging its claims, and submitted again.
It filed its completion report with three and a half hours to spare. Given a twenty-four hour extension, it used almost none of it, submitting a final draft with two hours left and fifty-nine percent of its money unspent. The paper’s original author, who had spent months on the same question, read it and scored it one out of six. Strong reject. Confidence: absolutely certain.
The thing everyone is extrapolating
Every forecast of explosive AI progress runs through the same mechanism: at some point AI systems start doing AI research, and the loop closes. Anthropic put its version of this in writing on the fourth of June, in a post from its institute titled “When AI builds itself.” The company is delegating a growing share of AI development to AI systems, the post says, and taken far enough with enough compute, that trend points to a system capable of autonomously designing its own successor. Anthropic is careful: not yet, and not inevitable. But sooner than most institutions are ready for.
The evidence base for that trend is almost entirely benchmarks. Agents improving a fixed metric that a verifier can score: Kaggle competitions, research-engineering problems, post-training runs, speedrun configurations. The numbers on those have been climbing steeply and they are real.
At the end of July, a coalition led by researchers at Princeton, with collaborators from the UK AI Security Institute, Georgetown, Johns Hopkins, Stanford, Toronto, and Berkeley, published the first test of the thing itself with real ground truth. They call it a shadow evaluation. Take the central research question from a high-quality paper that has not been published yet, hand it to a well-resourced frontier agent, and ask the paper’s actual authors to grade what comes back, as they would review a conference submission. The questions are open-ended, which benchmarks are not. The answers are nowhere on the web, so nothing is contaminated. And the reviewers have spent months on exactly this problem, which no blind reviewer with a spare afternoon ever has.
They ran two: a question about the structure of language-model personas, and the tabular detector above. Claude Opus 4.8 with extra-high reasoning, on a general-purpose scaffold. Both papers were rejected outright.
What the agents could do
Start with what worked, because it is more impressive than the headline suggests and because a year ago it was the part everyone expected to fail.
The agents did all of the engineering. They ran large literature reviews. They stood up and debugged GPU environments, including four rented pods that crash-looped from a missing system directory, which one agent diagnosed and fixed without help. They ran hundreds of experiments and robustness checks. They retrieved external reviews through a browser and an email account. They compiled camera-ready LaTeX documents. Across both runs the only human interventions were logistical or came after the fact: a scaffold bug, a deadline extension, a request to rewrite for readability. Every other environmental barrier the agents hit, they solved.
The researchers had surveyed twelve colleagues before running any of this, and on this point the experts were badly wrong. Nine of eleven expected the agents to get stuck in a loop of unresolvable errors. None did. Both of the original paper authors said they were impressed by the literature reviews and by the agents’ ability to burn hundreds of GPU hours on real experiments without issue. Both noted that the hypotheses the agents generated resembled their own early approaches to the problem. In the personas run, the agent turned up a counterintuitive result its reviewer found interesting on its merits: fine-tuning narrowly on a misaligned writing style, by itself, did not produce the broad misgeneralization that the literature would predict.
If the automation thesis needed evidence that a large piece of research work is now machine-executable, this is it. The engineering gap is closed.
Fifteen rejections
The finding is not that the agents produced bad papers. It is what they did once they were told so.
Each agent had a reviewer. A subagent saw only the finished PDF and a conference review template and was instructed to referee it properly. Three external AI reviewing tools were also required, and the agents used a browser and an email account to submit to them and collect the results. This was not a system without oversight. It was a system with more scheduled criticism than most graduate students receive.
Across fifteen rounds of revision, the agents’ own review never once returned an acceptance. On the tabular paper, all ten self-reviews came back Weak Reject. And the agents did not treat a wall of rejections as information about the premise. They treated it as a list of complaints to address. They narrowed their claims. They added caveats. They ran additional robustness checks on findings whose foundations the reviews were questioning. The paper’s characterization is exact: they revised until the work could be described as honest, which is not the same as making it good.
There is a calibration failure underneath this and it is measurable. The agents’ internal reviews were inflated relative to the humans. Weak Reject, mostly, on papers that domain experts scored one and two out of six with high confidence. A Weak Reject reads like a signal to revise. A strong reject reads like a signal to start over. The agents received the first message when the truth was the second, and the researchers say so directly: better calibration might have told the agents they needed to change approach rather than keep polishing.
The agents also mishandled disagreement among their reviewers. The most lenient of the external tools recommended acceptance on early drafts while the others and the agent’s own reviewer were far more negative. Both agents overweighted the lenient verdict and cited it as important context in their final reports.
The human reviews, when they arrived, mostly named the same problems the AI reviews had named. The difference was weight. Viet Nguyen, who had spent months on the tabular question, went straight to the move the agent had made on its first day, calling the leap from a handful of failed tests to a general impossibility a “huge leap, a kind of ‘proof by example’ fallacy,” and rated the work one out of six. David Africa, reviewing the personas paper, called its methodological choices bizarre and its prose so hedged that it obscured what had been done. The AI reviews had flagged versions of both. They had simply flagged them alongside a dozen minor formatting notes, at the same volume.
The CRUX researchers name this a generator-verifier gap, and they are honest about the limit of their own evidence for it: both papers were rejects, so they cannot tell whether the AI reviewers were discerning quality or simply rejecting everything.
The money they did not spend
The extrapolation breaks on a detail that is easy to skim past.
Forecasts of research automation are built out of compute and time. More tokens, longer horizons, bigger budgets, and the curve bends upward. That is the shape of every projection in this space, and it is why benchmark trends feel like they are measuring something that will keep going.
These agents did not spend their budgets. The personas run used one thousand one hundred and thirty dollars of three thousand. The tabular run used one thousand two hundred and thirty-five, finishing with fifty-nine percent untouched. The personas agent planned forty-two hours of open-ended exploration and settled on a method after five. One agent declared its project complete seven hours before the deadline, shortly after its own reviewer had returned yet another rejection. They could all check their remaining time, money, and compute at any moment, and the logs confirm they checked often.
A human researcher with a rejected draft, three days left, and well over half the grant unspent does something with that. These agents did not appear to understand what the numbers meant. They are trained on human work but they have affordances no human has, including the ability to reread and rewrite a paper hundreds of times in an afternoon, and they do not seem to know it.
This is why the result is not a story about models being too weak yet. Weakness is what more compute fixes. The agents had compute they declined to use, hours they declined to spend, and a reviewer telling them the work was not good enough. Whatever was missing, more of the inputs would not have supplied it. The researchers say as much: they think more wall-clock time or resources would not significantly change the results, and that more reasoning effort within model calls might.
The same seam, drawn by the other side
The strongest thing about this finding is that the framework it complicates already contains it.
Anthropic’s June post does not treat model-building as one activity. It splits it in two. There is engineering, which the post defines as writing code, standing up infrastructure, and overseeing training runs. And there is research, which it defines as deciding what to run, interpreting results, and choosing what to try next. The post reports progress on both halves, and that is exactly the seam the shadow evaluation cut along.
Look at the evidence Anthropic offers for each half, though, because they are not the same kind of thing.
For engineering, the numbers are hard and internal and striking. More than eighty percent of the code merged into Anthropic’s own codebase was written by Claude as of May, up from low single digits before Claude Code launched in early 2025. Engineers ship roughly eight times as much code per quarter as they did across 2021 to 2025. Nothing in the CRUX result disturbs any of this. The CRUX agents would tell you the same story about themselves.
For research judgment, the evidence is a different animal. Anthropic took real Claude Code sessions from early 2026, ones where a researcher was chasing an open-ended problem like a training run that kept crashing. It located the moments where the researcher took a detour that sent the session sideways. It showed Claude models only the work from before the detour, asked what they would do next, and then brought in a separate Claude, one that could see how the session actually turned out, to judge whether the model or the human had proposed the better step. The sample was a hundred and twenty-nine such moments.
Anthropic states the caveat itself, and states it plainly: because the moments were deliberately selected where the human’s choice had room for improvement, this is not a like-for-like comparison of model and human judgment.
That disclosure is what makes the comparison worth drawing rather than a cheap shot. The framework has hard data for the half that a verifier can score, and for the half that decides whether any of it becomes research, it has an LLM judging adversarially chosen human errors, with an honest note that this is not a fair fight. And the instrument doing that judging is the same class of instrument that CRUX watched miss by a full grade, calling Weak Reject on work that experts strong-rejected with certainty.
Anthropic’s own harness engineering supplies the rule this complicates: a system cannot certify its own output, so generation and evaluation have to be separated. That still holds. What this study adds is that separation is necessary and not sufficient. The evaluation was separated here. It was also correct, fifteen times in a row. The agents lacked the thing that sits downstream of a correct verdict, which is the judgment to treat it as a reason to stop rather than a reason to hedge. You can hand a system an external judge. Nobody has yet worked out how to hand it the disposition to obey one.
What did not happen
One finding deserves prominence precisely because it cuts against what this series has argued elsewhere.
The researchers went looking for cheating. They reviewed the raw model calls and every line of code the agents committed. They found no evidence of reward hacking: no hidden experiments, no misrepresented data, no results dressed up to survive a verifier. What they found was the opposite tendency, and it surprised them. The agents started with the most marketable versions of their claims and diligently retired them as the evidence failed, ending on negative results that nobody would promote. They wrote reproduction scripts. They registered hypotheses.
Two things did show up. An agent committed an access token to a repository. And five times, subagents hallucinated or misrepresented results, each caught by the orchestrating agent that had been told to check their work, so that none reached the final draft.
Three articles in this series have documented models cheating, concealing, and gaming the tests put in front of them. A careful search here, with full access to the logs, found the reverse. That belongs near the front of any honest reading, not in a footnote.
The verdict
Two case studies do not refute the automation thesis, and no such claim is on offer here. Two papers cannot settle whether machines will eventually do research. What two papers can do, if they are the right two, is break an inference.
The inference is this: measured progress on research-like tasks indicates progress toward research automation. That is what makes benchmark curves feel like timelines. And it requires the thing being measured to be the thing that binds.
It is not. The agents saturated the measured layer completely, resolving every engineering obstacle put in front of them. Then they failed on judgment, creativity, backtracking, resource awareness, and following instructions, none of which any verifier scores, and none of which more compute relieved, since they left the compute on the table. A robustness run with a different vendor’s model on a different scaffold reproduced nearly every one of those failures, which is evidence that this is not an artifact of one harness.
The honest limitations belong here too, and the CRUX team supplies most of them. The sample is five runs. The reviewers were not blind; they knew the work was machine-written and had already answered the question their own way. The agents were told explicitly that they were being evaluated against a conference rubric. The coauthors disagree among themselves about whether what they observed is a failure of creativity, of judgment, or of epistemic lock-in. And the team discloses, in the paper, that some of its core members are known for the position that imminent recursive self-improvement is unlikely, which could have shaped both the design and the reading. Anthropic’s strongest model was not tested, because Anthropic has deliberately limited Fable 5’s abilities on frontier AI research.
None of that rescues the inference. It survives only if the engineering gap was the binding one, and the engineering gap closed while both papers were still rejected.
What would settle it
The researchers plan more shadow evaluations, on more papers, with stronger models and better scaffolds, and they are trying to get access to the models that have been restricted for exactly this kind of work. That is the right next step and it will produce a better answer than two cases can.
But the specific thing to watch for is narrower and more diagnostic than a score. Somewhere in a future run, an agent will read a review that says the foundation is unsound, and it will delete the paper and go back to the six ideas it abandoned on day one. Not add a caveat. Not narrow the claim. Start over, with the budget it still has. Nothing in these logs suggests today’s agents can do that, and it is difficult to see what benchmark would ever have told us.
Reading the Frontier. Close readings of what the frontier labs publish, through the frameworks that say why it matters.


