Beyond Agents: LLMs as Invention Engines

Agents make existing work cheaper. The same models are already repurposing drugs, closing 80-year-old math problems and writing production algorithms — and the problems that qualify may be in your own stack.

Almost the entire conversation about large language models right now is about replacement. Agents that file the tickets, triage the inbox, write the boilerplate, run the support queue. The pitch is labour arbitrage, the metric is headcount avoided, and the demos are easy to film.

Meanwhile, quietly and with far less coverage, the same technology has been doing something considerably more interesting. It has been proposing drug candidates that survive contact with a laboratory. It has produced a result in theoretical physics that challenged a textbook assumption. It has designed proteins from scratch that bind to their targets better than the best human entries in an open competition. It has written algorithms that now run in production inside one of the largest computing estates on earth, recovering measurable compute every day.

These are not demos. Most are published and externally tested, and there are enough of them now to constitute a category. The reason you have read less about them is structural rather than sinister: an agent demo verifies itself in front of you, while an invention has to be checked by someone with a laboratory or a proof assistant, and that takes months. FutureHouse’s Robin went up as a preprint in May 2025 and appeared in Nature in May 2026. Sakana’s AI Scientist took roughly nineteen months to make the same journey. The news cycle that would have carried those results had moved three model releases on by the time they were confirmed.

So the argument of this piece is simple. If your organisation’s entire AI strategy is about doing existing work more cheaply, you are using the smaller half of the capability. The other half is acceleration of discovery — and unlike most things in this field, it is already evidenced, already funded, and available to anyone whose problems have the right shape. The last section of this piece is a test for whether yours do.

The evidence

A glaucoma drug, pointed at blindness

Give a system one line of instruction — dry age-related macular degeneration, a leading cause of irreversible blindness with few treatment options — and see what comes back.

What came back from Robin, FutureHouse’s multi-agent system, was a mechanism and a molecule. It proposed enhancing phagocytosis in retinal pigment epithelium cells, then identified ripasudil, a Rho kinase inhibitor already approved for glaucoma in Japan and which, to the authors’ knowledge, nobody had ever proposed for this disease. In vitro testing confirmed the effect. Robin then specified a follow-up RNA sequencing experiment to work out why, analysed the result itself, and surfaced ABCA1 — a lipid efflux pump whose lipid acceptor, apolipoprotein E, is a known genetic susceptibility factor for the disease — as a possible new target.

The paper, published in Nature, contains a sentence worth reading twice: all hypotheses, experimental directions, data analyses and data figures in its main text were produced by Robin. Not assisted by. Produced by. Humans ran the bench experiments; the intellectual framework came from the system.

Google DeepMind’s Co-Scientist, whose main paper appeared in Nature the same day as Robin’s, has produced the cleanest head-to-head we have. In a study published in Advanced Science, a Stanford researcher working on liver fibrosis picked two candidate drugs himself from the literature and asked the system for three. All five went through his lab’s testbed of live human liver cells. His two picks showed no benefit. Two of the system’s three blocked fibrosis and promoted the regeneration of liver cells. Among its candidates for acute myeloid leukaemia, all tested in the lab, was KIRA6, which killed leukaemia cells with an eighteen-fold margin over a non-cancerous control line — a promising in vitro therapeutic window.

Proteins designed from nothing, beating the field

De novo binder design means building a small protein from scratch that latches onto a chosen target. In August 2026, Anthropic reported that Claude, working from a written protocol but with no human input on any individual design decision, produced 354 confirmed binders against 14 of 15 targets from 1,320 designs, at hit rates of 22–35% depending on how the sessions were run, against what the company describes as a typical 10–15% for such campaigns.

The competitive number is the striking one. Adaptyv Bio runs public design competitions where human experts submit entries. Against a target called RBX1, only 9 of the 245 designs submitted to that competition bound (3.7%); the model hit 40% in single-target mode, and its top design outperformed the competition’s winning entry. Adaptyv Bio and Twist Bioscience produced and tested everything independently; the hit rates themselves are the company’s own figures.

Algorithms that are in production right now

This is the least romantic case and, for most readers, the most immediately relevant. Google DeepMind’s AlphaEvolve pairs Gemini models with automated evaluators that run and score every candidate program it writes. A scheduling heuristic it discovered has been in production since before the system was announced in May 2025, continuously recovering around 0.7% of Google’s fleet-wide compute. At that scale, a rounding error is a capital expenditure line.

It also found a way to multiply two 4×4 complex-valued matrices using 48 scalar multiplications — the first improvement on Strassen’s 1969 result in that setting in 56 years. The one-year update published in May 2026 reads like a list of unrelated specialist papers, all from the same engine pointed at different scoring functions: a 30% reduction in variant detection errors in PacBio’s sequencing model, a rise in feasible solutions for AC optimal power flow from 14% to 88%, quantum circuits with roughly ten times lower error than conventionally optimised baselines on Google’s Willow processor, and a 20% cut in write amplification inside Google’s Spanner database.

A physics result that was not supposed to exist

Standard textbook arguments hold that when one gluon has negative helicity and the rest positive, the corresponding tree-level scattering amplitude must be zero. In February 2026, a collaboration between the Institute for Advanced Study, Vanderbilt, Cambridge, Harvard and OpenAI published a preprint showing otherwise for a special “half-collinear” arrangement of momenta. The physicists had ground out cases up to six gluons by hand, producing expressions whose complexity grows superexponentially. GPT-5.2 Pro proposed a compact closed-form formula covering all of them, a scaffolded internal model spent about twelve hours producing a proof, and the human authors verified it. A physicist at UC Santa Barbara called it journal-level research advancing the frontiers of theoretical physics.

Problems that had been open for eighty years

Erdős left behind more than a thousand conjectures, catalogued at erdosproblems.com. According to a tracking page started by Terence Tao, AI tools helped move about a hundred of them into the solved column between October 2025 and February 2026.

The headline result is OpenAI’s disproof of the planar unit distance conjecture, posed in 1946, which nine mathematicians verified and rewrote in a companion paper published the same day. The construction borrowed machinery from algebraic number theory, a field built for entirely different purposes. What the mathematicians said about it is as instructive as the result. Most experts had agreed with Erdős and spent their effort trying to prove the conjecture, and the counterexample demanded a tedious construction with no encouraging signs along the way. A model has no career, no reputation and no fatigue, and so experiences the cost of an unpromising path completely differently from a person.

Which domains have actually explored this

Laid out together, the pattern in where progress has been fastest is not subtle, and it will matter when we get to the template.

DomainWhat checks the answerTime per checkWhere it has got to
Algorithms and codeAutomated evaluator running the candidateSecondsDeployed in production; measurable savings
MathematicsProof assistant plus expert reviewHours to monthsDozens of open problems closed; results entering journals
Theoretical physicsFormal proof and consistency checksDaysIsolated but genuine new results
Biology and drug discoveryWet lab or contract research organisationWeeksValidated candidates; nothing approved yet
Materials and chemistryRobotic synthesis and characterisationWeeks to monthsHeavily funded; the loop is still being built

At the frontier sits the most spectacular and least settled claim of the year. On 8 September 2026, OpenAI announced that an internal system had shown the Navier–Stokes equations can develop a singularity in finite time — which would resolve one of the seven Millennium Prize Problems — with a 166-page manuscript and a Lean formalisation, produced by roughly 10,000 concurrent agents over about 88 hours at a cost it puts in the millions of dollars. Two weeks on, the picture is more interesting than the headline. The Clay Mathematics Institute says its evaluation will be deliberately unhurried, and a dispute over credit remains open with two mathematicians — one of them using an internal Anthropic model — who had reached a related result on the Euler equations a day earlier. More telling, the proof works by exploiting an option in Clay’s official problem statement that allows a precisely contrived external force, and three mathematicians have since posted a proof that the method can never be extended to the force-free version most experts actually care about. As one Chicago mathematician put it to Scientific American, the Clay problem may be settled while the main problem is not. I include it as a marker of where the frontier is — and, as we will see, as a lesson about verifiers.

The invention engines being built

What was a research curiosity two years ago is now a funded product category, and the capital tells you where practitioners think the value sits.

In biology, FutureHouse spun out Edison Scientific to commercialise Kosmos, a system whose single run of up to twelve hours reads roughly 1,500 papers and executes some 42,000 lines of code. Collaborators estimated that one run equalled about six months of their own research time, and independent scientists judged 79.4% of its statements accurate. The claim I find most significant is quieter than the six-month figure: Edison reports that the number of valuable findings grows roughly linearly with run length, which is effectively an inference-time scaling law for research — spend more compute, get proportionally more science. Kosmos has moved into biopharma, including a partnership with Incyte.

In mathematics, the engines are built around Lean, the proof assistant that can mechanically check whether a proof holds. Harmonic’s Aristotle formalised the negation of the Erdős unit distance conjecture and has topped the Lean formalisation leaderboard; it raised $120M at a $1.45B valuation, with Nvidia among the backers. Axiom Math raised $200M at $1.6B in March, pitching “verified AI” for software as much as for mathematics, and in August its prover became the first to formally verify the “246 theorem” on bounded prime gaps. Math Inc’s Gauss formalised Viazovska’s Fields Medal-winning sphere-packing proof in 8 and 24 dimensions in roughly 200,000 lines of Lean. Roughly three billion dollars of valuation now sits on machinery whose only job is checking whether mathematical claims are true.

In the physical sciences, the engines are robots. Periodic Labs, founded by former OpenAI and DeepMind researchers, launched in 2025 with a $300M seed round for autonomous materials discovery and has reportedly held talks at a valuation around $7B. Lila Sciences has raised $550M at a valuation above $1.3B for autonomous laboratories spanning chemistry, life sciences and materials, and is reportedly in talks at a far higher figure. This is the least mature of the three: as of late 2025, Lila’s labs were largely empty and Periodic was beginning with manual synthesis guided by AI predictions, and neither has publicly demonstrated a fully autonomous discovery-to-synthesis loop at scale.

In algorithms, AlphaEvolve moved from internal tool to Google Cloud offering in May 2026. Early partners report gains on their own figures — Klarna doubling transformer training speed, Schrödinger speeding up molecular force-field models about fourfold — though independent enterprise results do not yet exist.

Notice what nobody is funding: better generators for science. Those are rented from the frontier labs at commodity prices. The money is going into machinery that checks answers — proof assistants, robotic laboratories, assay capacity, evaluation infrastructure. That is a strong signal about where the scarce resource actually is, and it anticipates the test at the end of this piece.

What is actually happening when we say a model “invented” something

Being precise here is not a way of taking the results back. It is the difference between knowing when this approach will work for you and hoping.

Every system above has the same three parts. A generator proposes candidates — a proof sketch, a molecule, a protein sequence, a scheduling heuristic. A verifier decides whether a candidate is any good: a proof assistant, a test harness, a physical assay, an expert with a pencil. A search loop feeds the verifier’s judgement back into the next round. The LLM is the generator. It is not, by itself, the inventor, and the systems that produce durable results are disciplined about this — AlphaEvolve discards a plausible-looking but wrong program by executing it, not by judging it.

Within that loop, four capabilities do most of the work, and none of them is creativity in the mystical sense:

  • Recombination across a literature no individual can hold. Ripasudil was not invented; it was connected. The invention was the link between a glaucoma drug and a retinal disease, which lay latent in a body of work too large for any person to have read. The unit distance counterexample did the same across fields, carrying number theory into geometry.
  • Generalisation from worked examples. The gluon formula came from spotting a pattern across hand-computed cases and proposing the closed form nobody had found.
  • Patience without ego. Grinding through an unpromising construction for as long as it takes, with no reputational cost to a dead end.
  • Orchestration. The protein campaign invented no new design method; the model drove existing open-source tools and made the specialist judgement calls that normally gate their use.

Once you see the shape, the limits follow from it rather than arriving as a list of complaints:

The generator is wrong most of the time, and that is fine if checking is cheap. In DeepMind’s own sweep of 700 open Erdős problems, of the 200 candidate solutions humans graded definitively, only 13 — 6.5% — were meaningfully correct. Roughly one statement in five from Kosmos is inaccurate. And in the Robin paper itself, the system’s automated analysis reported that ripasudil boosted phagocytosis 7.5-fold; human re-analysis of the same data put it at 1.75-fold. The drug still worked. The generator’s own read of the result was badly inflated. This is a search process, not an oracle, and its economics only work when the verifier can absorb the volume. Hit rate per candidate, not raw output, is the number to ask any vendor for.

Rediscovery looks exactly like discovery from the outside. In October 2025 an OpenAI executive posted that GPT-5 had solved ten previously unsolved Erdős problems. The database maintainer called this a dramatic misrepresentation: “open” on his site meant only that he personally was unaware of a paper solving it, and what the model had found were existing references. The post was deleted. Finding a forgotten solution buried in a decades-old paper is genuinely valuable. But it is retrieval. Of the 13 problems DeepMind’s sweep resolved, nine turned out to have solutions already in the literature, and its authors warn of “subconscious plagiarism” — a model reproducing something it absorbed in training without knowing it. Conflating retrieval with invention is what makes serious readers discount the whole category.

The verifier is the trust anchor, so it has to be sound. Formal verification is the strongest checking we have and still not absolute: a recent demonstration showed how a bug in the method could be exploited to accept a false, AI-generated proof. Subtler, and more common, a verifier checks the statement you gave it, not the question you meant. Of the 63 Erdős answers DeepMind’s team found technically correct, 50 solved a misreading of the problem, often trivially. The Navier–Stokes proof satisfies the letter of Clay’s statement through an option most experts consider beside the point. A generator searching hard against a checker will find the gaps in how the target was written down. If you build one of these loops, the evaluator and the problem statement are the components that deserve adversarial review, because everything downstream inherits their errors.

Much of the best work is not reproducible by outsiders. Several headline results above came from unreleased models; DeepMind’s FullProof is private; Google has declined to open-source Co-Scientist’s code and weights; and every verified AlphaEvolve result to date is internal to Google or involves a Google partner. Treat vendor-reported numbers accordingly, and prefer results with external validation attached.

A template: does your problem qualify?

Every success above shares the same structure, and it is a structure you can test for. Six questions, in order of how much they matter:

  1. Can you score a candidate answer without human judgement? A benchmark, a test suite, a simulation, an assay, a proof checker. If scoring requires a meeting, the loop will not run. And check that the score measures what you actually want, because the search will find every place it does not. This is the question that decides it; the rest are refinements.
  2. How long does one check take, and what does it cost? Seconds means thousands of iterations. Weeks means dozens, and you will need a cheap proxy check in front of the expensive real one — which is exactly what in-silico screening ahead of wet-lab validation is.
  3. Is the search space larger than a team could explore by hand, but small enough that good answers exist in it? Scheduling heuristics and binder sequences qualify. “Design a better product strategy” does not.
  4. Is a better answer worth anything? 0.7% of a fleet is worth a great deal at Google’s scale and nothing at ten machines. Multiply the plausible improvement by the volume it runs at before you start.
  5. Can you tell novelty from retrieval? You need a way to know whether a proposal is genuinely new or already in your own archive, your vendor’s documentation, or the literature. Either is useful; confusing them will burn your credibility internally.
  6. Can you accept the answer safely if it is strange? These systems produce solutions no human would have written, sometimes for reasons nobody can articulate. You need a rollout path — canary, secondary safety metric, rollback — that does not depend on a reviewer understanding why it works.

Applied to ordinary infrastructure, the qualifying problems are more common than they look: scheduling and placement, capacity allocation, cache and index strategy, query planning, compression and encoding choices, parameter and configuration tuning, test suite minimisation, routing policy. Each has an objective you already measure, a candidate space too large to hand-explore, and a verifier you probably already own in the form of a benchmark. The disqualifying ones are just as recognisable: anything where success is contested, where the cost of a bad candidate reaching production is catastrophic and unmonitorable, or where you would need a person to read every proposal.

What to do with this

The instinct when starting one of these projects is to pick the model first and ask what it can do. The evidence says the opposite: start with what you can score automatically, spend the first month building the evaluator rather than the agent, and treat the generator as interchangeable — because it is, and it will be replaced twice before the project matures.

For anyone allocating capital rather than engineering time, the same logic points somewhere less crowded than model access. Generators are a competitive market with falling prices. Validation capacity — automated laboratories, formalisation infrastructure, proprietary experimental data, instrumentation — is physical, slow to build and hard to copy, and it is what every serious discovery effort is currently short of.

And there is a reasonable counter-argument that deserves stating. If model quality keeps rising, the generator may eventually propose so few wrong answers that verification stops being the constraint at all — that 6.5% becomes 60%, and the careful machinery matters less than the raw capability. Everything above describes the present distribution of results rather than a law of nature. The number that will tell you which world you are in is the hit rate, not the headline count.

Either way, the strategic point stands. Deploying agents to do cheaply what your people already do is a cost programme with a floor: you cannot save more than you currently spend. Pointing the same models at problems nobody has solved yet has no such ceiling. The organisations that work that out early will be the ones that noticed the second half of the capability while everyone else was counting the headcount they saved.