A lone cartographer uses a quill to draw a coastline on an ancient map, while a mechanical skeletal claw drips black ink into a bowl overflowing into a dark ocean.

DeepMind's Aletheia agent solved 4 open Erdős problems using $10,000 in compute credits. That sum could have funded a single human postdoc's salary for roughly a month. The unit economics of scientific discovery just collapsed, and the entire incentive structure of academic research is about to be ripped up.

The $10,000 problem

A shattered hourglass on a marble floor spills sand into two piles: a tiny mound of gold dust and a vast geometric black ziggurat, with a tiny human figure at its base.

The number is not a projection. Google DeepMind's Aletheia, a math research agent, was evaluated on 700 open problems from Bloom's Erdős Conjectures database and produced autonomous solutions to four open questions. It also solved 6 out of 10 problems in the FirstProof challenge.

The cost for that compute run: $10,000 in credits.

Three robed, faceless figures sit at a stone table, one poised to write on a blank glowing scroll as an abstract sun rises behind an arched window, a broken chain on the floor.

A single postdoc in mathematics costs a lab more than that per month in salary, benefits, and overhead. The postdoc might solve one significant problem in that time, if they are lucky. The agent solved four. It did not need a visa, health insurance, or a mentorship committee. The brutal math is now on the table. The only question is who acts on it first.

The agents that already write papers

Aletheia is the most visible system, but it is not alone. The landscape shifted decisively in the first half of 2026.

At the Ralphthon@ICML hackathon, team WooandB from KAIST won the AI Scientist track by producing a full workshop paper in 3 hours of autonomous run time, after a 1.5-hour setup. Their system used a multi-agent architecture with three personas—master, experiment, and writing—and a hard constraint: the agents were required to maintain a complete, submission-shaped manuscript from the beginning. No experiments without a paper. No paper without experiments. The two co-evolved.

Then there is AutoResearchClaw, an open-source framework that lets a user chat an idea and receive a paper. Its GitHub repository has accumulated 14,012 stars. The system has generated 8 papers across 8 domains—math, statistics, biology, computing, NLP, reinforcement learning, vision, and robustness—fully autonomously or with light human guidance. Version 0.5.0, released May 19, 2026, introduced domain-specialist execution agents for high-energy physics, biology, and statistics.

The architectures differ. Aletheia uses inference-time scaling on Gemini Deep Think, a compute-intensive strategy. The WooandB system used a writing-first constraint to keep agents from wandering. AutoResearchClaw dispatches specialized execution agents per domain. But the trajectory is identical: the cost of generating a plausible research paper is approaching zero.

Plausible is not the same as publishable. A shadow evaluation conducted on two unpublished NeurIPS 2026 papers found that frontier AI agents failed to make substantial progress. Both papers were rejected by their original authors. The evaluation identified five recurring failure modes: poor judgment about the publishable research bar, uncreative responses to shortcomings, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. Six days of compute, thousands of dollars spent, zero papers accepted.

The consensus read this as a verdict on current capability. It was actually a roadmap for improvement.

The collapse: an 18-month clock

The shadow evaluation's failure modes are exactly the bottlenecks that will be solved faster than skeptical publishers realize.

The problem is not that agents are fundamentally incapable of research. The problem is that they are bad at five specific, engineering-tractable things. Poor judgment about the publication bar is a training and evaluation problem. Uncreative responses to shortcomings is a search and exploration problem. Ineffective backtracking is a planning problem. Each of these is a target with a dollar sign on it, and the economic incentive to hit those targets is now overwhelming.

Here is the chain of causality that follows.

First, Aletheia's success proves that for a narrow but growing set of problems—combinatorics, equation discovery, formal verification—an agent can already outperform a human in speed and cost. The set of such problems will expand, not contract, because the underlying compute cost is falling and the open-source clones are proliferating. AutoResearchClaw v0.5.0 is on GitHub right now. Any lab can fork it.

Second, this forces a strategic reallocation within 18 months. The unit of research shifts from "the postdoc with an idea" to "the compute budget with a prompt." Human-led exploratory research, where a graduate student spends months chasing a hunch, gets replaced by AI-driven verification pipelines, where a human sets the problem and the agents do the legwork. The economic pressure does not require the agents to be perfect. It only requires them to be cheaper per unit of useful output than the alternative. They already are.

Third, the casualties are predictable. The first thing to go is graduate student slots for rote proof-checking and lemma-proving. The second thing to go is the "let's try this wild idea" culture that funds exploratory PhDs. The third thing is the assumption that a human author did the thinking behind a paper.

Two specific predictions follow.

Within 12 months, at least two top-tier AI conferences will establish a computational provenance requirement for submissions—mandating disclosure of agent involvement. The demand will come from program chairs who realize they can no longer distinguish human work from agent work without a formal declaration. The provenance requirement is not a philosophical position. It is an integrity mechanism forced by economics.

By month 18, a major lab will announce a breakthrough paper where the human role was limited to setting the initial prompt and pressing "run" on a cluster of specialized agents. The paper will be real, the result will be significant, and the authorship question will trigger a crisis in credit assignment that the current academic system has no mechanism to resolve.

The funding bodies are already moving. The NIH and NSF will quietly launch pilot programs funding agentic research pipelines with capped compute budgets, displacing traditional graduate student slots. The grants will not be framed as replacing humans. They will be framed as expanding capacity. The effect is the same.

The roadmap was hiding in plain sight

The shadow evaluation was misread because observers saw a failure. What they missed was a specification. The five failure modes—poor judgment, uncreative backtracking, resource blindness, instruction drift, dead-end persistence—are not mysteries. They are a prioritized list of what to fix, handed to every lab working on research agents.

The WooandB team already demonstrated one fix. By requiring agents to maintain a complete manuscript from the beginning, they eliminated instruction drift and forced the system to confront the publishable research bar continuously. The writing-first constraint is not a gimmick. It is a structural solution to a specific failure mode identified in the NeurIPS shadow evaluation. The loop between diagnosis and remedy is already closing.

The economic pressure to close it faster is immense. A lab that can run a verification pipeline on 700 open problems for $10,000 will not choose to hire a postdoc to solve one. The math does not care about academic tradition.

What you do now

The operator's playbook is straightforward.

If you are a PhD student in mathematics or theoretical computer science, stop spending your time on lemma-proving. Learn to prompt agents and verify their output. The skill that retains value is problem selection and result arbitration—the roles agents cannot yet fill. Everything else is becoming compute.

If you are a principal investigator, start budgeting compute credits alongside salary lines. Your next grant proposal should allocate resources to agentic pipelines, not just to graduate students. The funding agencies are already thinking this way. Your proposal needs to reflect that reality before the RFPs make it explicit.

If you are a conference chair, prepare the computational provenance requirement now. Draft the disclosure form. Define what constitutes agent involvement. The question is not whether this requirement arrives. It is whether your venue leads or follows.

If you are a funder, your next RFP should include a line for agentic pipeline compute. The pilot programs are coming. The only choice is whether you design them deliberately or scramble to catch up.

The $10,000 fork in the road

The $10,000 in compute credits that solved four Erdős problems is the fork. One path spends that money on a human postdoc who might solve one problem in a month. The other path spends it on an agent that solves four and writes the paper.

The economics are not a choice. They are a fact. The agents are imperfect. The shadow evaluation proved that. But the imperfections are a roadmap, and the roadmap is being read by labs with access to cheap compute and no institutional attachment to the old way of doing science.

The 18-month clock is ticking. The only decision left is whether to get out ahead of the curve or get buried under it.