
A Google DeepMind agent just solved 4 open problems from the Erdős conjecture database without human help.
This is not a contest puzzle. Not a curated benchmark. These are problems Paul Erdős left unsolved, sitting in a database of 700 open conjectures, waiting for a professional mathematician to crack them. Aletheia, an autonomous mathematics research agent developed by Google DeepMind, evaluated all 700, identified 4 it could solve, and produced verifiable proofs. No human in the loop. No hybrid methodology with expert checkers reviewing output before submission. End-to-end autonomous.

Six weeks earlier, on February 5, 2026, a separate challenge called FirstProof released 10 research-level lemmas. The deadline was February 13. Aletheia submitted solutions for all 10. A panel of expert mathematicians assessed the output. Six of the 10 were judged correct by majority vote. On Problem 8, the experts split: 5 of 7 rated it correct. The agent ran each problem twice and took the best output. That was enough.
This is the threshold. AI has moved from pattern-matching on competition math to generating novel, verifiable proofs in professional research. The paper is on arXiv. The results are reproducible. The implications for mathematics departments, funding agencies, and anyone who builds research tools will land within 18 months.
The 4 problems that broke the seal
The Erdős case study is the sharper signal. In earlier work, a semi-autonomous system using Gemini 1 addressed 13 problems marked "Open" in Bloom's Erdős Conjectures database. Four of those produced seemingly novel solutions. Nine turned out to have been solved previously in obscure literature, and the system identified those existing solutions. The authors noted the "Open" status likely reflected obscurity rather than difficulty.
Aletheia's autonomous run is different. It used Gemini 3 Deep Think, an iterative generation-verification-revise loop, and natural language output. The agent evaluated 700 open problems, selected its targets, and produced publication-grade papers. One paper required zero human intervention. This is not a solver bolted onto a search engine. It is a research agent that decides what to work on, attempts a proof, checks its own work, and revises until the proof holds.
The distinction matters because pure math has always been the hard case for AI. Competition problems have known solution templates. Research problems require generating structure where none is visible. Aletheia crossed that line.
The verification loop that makes it real
Here is the mechanism. Aletheia is powered by Gemini Deep Think. It operates in a loop: generate a candidate solution, verify it against formal criteria, revise if it fails, repeat. The output is natural language, not code or a formal proof assistant trace. That means human mathematicians can read the proofs directly. It also means verification is not automatic. On FirstProof, DeepMind used a best-of-2 run per problem and submitted results to a panel of human experts.
Problem 8 exposed the friction. Five of seven experts said correct. Two did not. That disagreement is normal in peer review. It also tells you the system is operating at the boundary where judgment calls matter, not just mechanical verification.
This verification pipeline is the part that will commoditize. The paper proposes "human-AI interaction cards" for transparent documentation of AI-assisted results. If DeepMind open-sources the verification pipeline, anyone can run autonomous proof-checking. The cost of testing a hypothesis in pure math drops toward zero. A graduate student with a laptop could evaluate hundreds of open conjectures overnight and wake up to a ranked list of which ones yielded to the agent.
The 18-month reallocation
Here is what changes and who feels it first.
Mathematics departments. Within 12 to 24 months, at least one major department will establish a dedicated AI proofs lab. The pressure is straightforward: if your competitors can test 700 conjectures in a single run, you cannot afford to ignore the tool. The first peer-reviewed paper with an AI agent listed as co-author will appear in a top-tier journal, likely the Annals of Mathematics. Editors are already grappling with authorship policies for autonomous systems. Aletheia's publication-grade output forces the conversation.
Funding agencies. The National Science Foundation and its counterparts allocate grants based on human investigator track records. When an autonomous agent can produce novel proofs at scale, the unit economics of discovery invert. A grant that funds one postdoc for a year could fund thousands of agent-hours of exploration. Within 18 months, expect funding calls that explicitly require or reward AI-augmented methodology. Agencies that delay will see their best applicants migrate to institutions that built the infrastructure early.
The convergence with Karpathy's autoresearch. On March 6, 2026, Andrej Karpathy released the autoresearch repository. It gives an AI agent a small LLM training setup and lets it experiment autonomously overnight, modifying code and training for 5 minutes per iteration. The repository has 93,229 stars and 13,258 forks. That is not curiosity. That is latent demand for autonomous research tooling.
Aletheia and autoresearch target different domains, but the architecture converges. Both use an agent loop: generate, test, revise. Both aim to commoditize the iteration cost. A cottage industry of "AI research assistants" will emerge from this intersection. The cost of hypothesis testing in pure math will drop by 90 percent, not because the models get smarter, but because the verification pipeline becomes cheap, standard, and open.
The losers. Mathematicians who dismiss AI as irrelevant to their field. Not because they are wrong about the current limitations. Because the cost curve has bent. When hypothesis testing costs drop by an order of magnitude, the researcher who refuses the tool is competing against peers who test 10 times as many ideas. That is not a philosophical disagreement. It is an economic mismatch.
The winners. Institutions that embed these agents into graduate training programs. A PhD student who learns to steer an autonomous proof agent will outproduce one who does not. The gap compounds. Departments that build AI proofs labs now will attract the best students and the most grant money. The first-mover advantage in pure math is usually small because the pace of discovery is slow. Autonomous agents change that equation.
What operators should do now
If you run a mathematics department, allocate a seminar room, a compute budget, and a postdoc to explore Aletheia's pipeline the week it is open-sourced. Do not wait for a consensus paper. The consensus will arrive after the first-mover window closes.
If you are a graduate student, learn to work with these tools immediately. The skill is not using the agent. The skill is asking the right questions, interpreting the output, and spotting where the agent is confident but wrong. Problem 8 on FirstProof is the template: 5 experts said correct, 2 did not. The researcher who can adjudicate that split will be indispensable.
If you are a funder, write the RFP now. Autonomous proof agents are not a future capability. They are a present one. The first agency to fund AI-augmented pure math at scale will own the narrative and attract the best proposals.
The cost of hypothesis testing in pure math is about to drop by 90 percent. That is not a prediction about model capability. It is a statement about the verification pipeline. Once the pipeline is open, the bottleneck shifts from "can we solve this" to "which problems should we aim at." Strategy, not compute, becomes the scarce resource.
The first domino
The 4 Erdős problems were the first domino. Not because they were the hardest. Because they were autonomous. No human in the loop. No hybrid methodology. Just an agent, a database of 700 conjectures, and a verification loop that produced proofs a mathematician can read and judge.
The cascade is now inevitable. Aletheia's paper is on arXiv. The FirstProof results are published. The autoresearch repository has 93,000 stars and counting. The economics of discovery have shifted. The only remaining question is who moves first.
The proof is in the proof.