top of page
Search

Towards Autonomous Science

3 hours ago
10 min read

AI can now iterate its way through problems with verifiable answers; the much harder task for true autonomous science is figuring out discovery, real-world experiments, and knowing what to look for to begin with.



Written by: George Patin | Research Contributor, London Venture Capital Network


In February, OpenAI's GPT-5 was given control of Ginkgo Bioworks' robotic laboratory and asked to make a protein more cheaply. Over six rounds it designed and ran 36,000 reaction compositions on more than 580 plates, generated about 150,000 data points, and brought the cost of cell-free protein synthesis from $698 a gram to $422. Humans prepared the reagents and loaded the plates. Later in September, as just about everyone has heard by now, OpenAI reported that an unreleased model, coordinating roughly 10,000 agents, had proved finite-time blowup for the forced three-dimensional Navier–Stokes equations, and had formalised the proof in Lean so that a compiler could check it.


Both get labeled as autonomous science. The two, however, are entirely different challenges.


One is essentially a massive iterative engine: a swarm searching within a defined space with answers that are computationally verifiable. The other requires the software to bridge to hardware -- actually form hypotheses, design and run experiments -- something that, for now, machines cannot quite do all by themselves.


Both are frontier science, and yet this distinction shows precisely which parts of autonomous science are likely already here, and which will not arrive for quite some time.


Why This Became Possible Now


The ambition is old. Adam, a robot scientist that generated and tested hypotheses about yeast gene function, was published in 2009. We have come some ways since then. The new part is that the loop now closes around a general-purpose model, with a crop of fresh startups now seeking to build out this capability and apply it to increasingly complex research.


Beyond the meteoric progress in frontier-model capabilities, three things have shifted.


The first is task horizon: agentic systems can now pursue goals across hours, days and, most recently, weeks. Berkeley's A-Lab and the Coscientist system showed in 2023 that a model could plan a synthesis; what has improved since is reliability over long sequences.


The second is sheer scale: you can buy parallelism, albeit at a steep price. Where results can be checked mechanically, you can run attempts by the thousand — exactly what the Navier–Stokes run did. Combined with longer task horizons, this means that, with enough money, you can set a very stubborn autonomous swarm on a problem until it cracks it. “AI for Science” doesn't really mean GPT-6 going off and winning a Nobel Prize — I would shelve that until GPT-11 at current rates. The discoveries getting the most attention were achieved by

orchestrated swarms of agents: iterating, setting hypotheses, invalidating them and trying again until something clicks.


The third is that the physical layer is slowly being bridged to the software: orchestration platforms that make multi-vendor instruments programmable, cloud labs that rent capacity by the protocol and, since August, a proposed common interface for laboratory hardware. For now, this is where AI is least reliable. Few foundation models for lab biology or general-purpose robotics are good enough to run a physical lab unaided, and few labs have been rebuilt to let them.


Taken together, it is not intelligence that is scarce but data and connectivity. Autonomous discovery is already showing real promise in domains like mathematics, where much of the work runs on pure computation. The barrier is interfacing with physical equipment, which limits the ability to verify hypotheses experimentally in domains such as biology or materials.


This, then, is progress to watch.


Autonomous Labs ≠ Autonomous Labs


“Autonomous lab” — and, for that matter, “AI for Science” — currently describes at least four businesses with different mechanics and different economics. What separates them is not how capable the model is, but what it costs to find out that the model is wrong.


In formal mathematics, verification is reasonably straightforward: a proof compiles in Lean or it does not. In computational chemistry and protein design, the check is a simulation — fast and cheap, but often circular, since one model is being validated against another. In a wet lab, it is instrument-hours, reagents and a couple of sleepless nights. In open discovery — the claim that something is new, works and matters — the checker is the rest of the scientific community, including independent replication, which takes months or years and cannot be bought with compute.


Autonomy tracks that cost more closely than it tracks model quality. That is why swarms are closing theory problems this year while the most advanced physical systems are still optimising known processes.



So a company calling itself an autonomous lab may be running an agent swarm on theory, a co-scientist over literature and data, a robotic workcell on a known process, or a full hypothesis-to-experiment loop over matter. The first two are now very real. The third is getting there. The fourth is what much of the valuation rests on, and it is likely still some years away.


The Landscape: Four Layers, Four Kinds of Cheque


Each layer comes with its own capital profile alongside a distinct lens on the AI for Science thesis.



Around $1.6bn of disclosed private capital has gone into the companies named here since the start of 2025 — a tally of announced rounds, not a market figure — with another $2.5bn reported but unconfirmed. Lila Sciences has raised a $200m seed and a $350m Series A, the latter closing in October 2025, and has been reported in talks to raise about $2bn at an $8.5bn pre-money valuation, anchored by CalPERS and Nvidia's NVentures. Periodic Labs, founded in May 2025 by Liam Fedus, formerly of OpenAI, and Ekin Doğuş Çubuk, formerly of Google DeepMind, raised a $300m seed at $1.3bn and has been reported in advanced talks at around $7.5bn — almost six times the seed price, eight months later.



Two features stand out in the capital allocation. Nvidia's venture arm sits on three of the six cap tables above, Nvidia turns up again at Orbital and Ineffable, and it is an industrial partner in CuspAI's materials coalition. And the cheques increasingly come from outside venture: a US public pension fund, sovereign vehicles, and strategics including Danaher, RTX, AMD and Eni.


Public money behaves differently. The US Genesis Mission, set up by executive order in November 2025, announced its first wave on 22 July: 278 awards across 342 institutions, from what the Department of Energy calls the largest response in its history. The White House headline was more than $5bn in federal commitments; the awards themselves draw on a $293m pool, with the largest single project at $60m over three years. The UK's AI for Science Strategy commits up to £137m. Within it, the Sovereign AI Unit has opened a call for autonomous-lab infrastructure with awards reported at £1–9m, and ARIA doubled its AI Scientists programme to £6m across 12 projects chosen from 245 proposals. Horizon Europe has set aside €33m for automated laboratories within a €100m AI-in-science pilot.


The split is clear enough. Public money is buying the commons — shared instruments, datasets, standards, national compute. Private money is buying proprietary loops.


Where the Closed Loop Already Delivers


The Ginkgo campaign is the cleanest public example: a model designing experiments, reading results and prioritising reagents over six rounds, and arriving at compositions materially cheaper than the prior state of the art. It is a preprint rather than a peer-reviewed result, and it covers one protein under specific conditions, but it points firmly in one direction.


What matters is the shape of the problem: a bounded parameter space, an agreed metric (cost per gram), and a measurement that happens in-line, quickly, and settles the question. Give an autonomous system those three properties and it will usually beat a human team — and then keep going, since AI does not exactly sleep.


For now, this is the template being replicated most. Practitioners describe a five-rung autonomy ladder, from scripted execution up to workflows that transfer between instruments, sites and operators. By the assessment of the venture and trade analysts who track the category, most currently deployed systems sit somewhere in the middle of it.



Where Open Discovery Breaks


Right now, open discovery fails in three distinct ways.


One caveat first, though. Much of the field is made up of private companies. It is in principle possible that Lila, Periodic or other labs have already advanced further and are holding results back for competitive advantage; OpenAI is probably not the only player with unreleased, internally capable models. The disclosure lag is short: I would expect any major discovery to be made public quickly, and would place the gap between the “true” and the publicly visible frontier at about six months. The verification lag is likely longer: independent, peer-reviewed results can take years to arrive.


This is to say that, altogether, the frontier is likely a little over the horizon from what is publicly visible. Caveats notwithstanding:


1 · Judgement


The best-evaluated systems are strong at the analytical half of science and measurably weaker at the interpretive half. Edison's Kosmos is one of the few with a published external evaluation: scientists assessed 102 statements from three representative reports and found 79.4% of them accurate. By type, that was 85.5% for data analysis, 82.1% for literature review and 57.9% for what the paper calls synthesis statements — the ones that draw findings into a conclusion. Extrapolate beyond a single report, and the ability of current systems to combine conclusions, null results and promising directions across an entire domain is likely weaker still.


Novelty is harder again. Berkeley's A-Lab reported making 36 of 57 target compounds over 17 days and 353 experiments. A critique by chemists at Princeton and UCL argued that the materials were already in the crystallographic databases; Nature issued a correction in January, but the dispute is not fully settled. The hardest public benchmarks point at a similar barrier. Epoch's FrontierMath Erdős set, released on 1 September, contains 68 significant unsolved Erdős problems formalised in Lean. One model has solved two; every other model tested scored zero.


Simply put, there is still a major gap between extrapolating from prior research and autonomously setting a direction that leads to novel discovery. The current tech isn't quite there yet.



2 · Coordination and Navier–Stokes


The Navier–Stokes case is perhaps the best illustration of where large-scale autonomous science currently stands.


The mechanics first. Roughly 10,000 agents, working concurrently, reached a solution in about 88 hours, producing on the order of 130 billion output tokens across 2.7 million messages; formalising and verifying the proof in Lean took another 17 hours. Outside estimates of the compute bill run from several million dollars to tens of millions.


Two elements stand out. First, the result concerns the forced equations: a specifically constructed external force drives the blow-up. Mathematicians quoted by Scientific American regard that as technically satisfying one of the options in the Clay Mathematics Institute's 2000 problem statement. But the unforced case remains open, with several mathematicians arguing the method cannot reach it. The swarm cracked the version of the problem it was capable of verifying.


Second, there is the provenance dispute. Two mathematicians posted related work twelve hours earlier, and one alleged that OpenAI moved after learning of their approach. OpenAI denies seeing the work before publication but concedes it cannot rule out that de-identified data from product usage improved its models. Whoever is right, the dispute is a harbinger of a much larger problem: there are no settled norms for credit, priority or provenance in AI-driven science, and science is a fundamentally collaborative effort. Nor is it yet clear what peer review looks like when the work arrives this fast.



A constructive counter-example is a collaboration between Google DeepMind and academic mathematicians a year earlier: it used physics-informed neural networks to find new families of unstable singularities in simpler fluid equations, at precision orders of magnitude beyond previous attempts. In essence, DeepMind gave mathematicians a tool and left the choice of where to point it to them. This is not full autonomy, but it is a reasonable guess at where autonomous science is heading: autonomy of execution, iteration and validation, but perhaps not autonomy of judgement.


3 · The Physical Layer


Pointing AI at purely computational challenges is one thing: to grossly oversimplify, the process is not that different, structurally, from iterating a company-wide go-to-market strategy. Linking up with physical infrastructure is a different matter entirely.


To return to Ginkgo: humans still handled the reagent preparation, the plate loading and the rest of the bench work. Most laboratory instruments still lack a usable programmatic interface, and most systems sold as self-driving coordinate one instrument rather than an array of them. This is where many startups are now focused. Three orchestration platforms launched within weeks of each other, and in August Anthropic proposed a Model Hardware Standard, an interface that lets agents discover and drive instruments. It was developed with HHMI's Janelia campus, counts Automata, Danaher, QIAGEN, Tecan and Universal Robots among its partners, and comes with a promised open-source release.


Getting the software–hardware bridge working is only half the story, however. Throughput is set by the slowest measurement, not the fastest robot: fast synthesis paired with slow characterisation leaves most of the parameter space unexplored. Data provenance is patchy — calibration state, reagent lots and ambient conditions are rarely recorded well enough for a result to transfer between sites. That is why this year's literature keeps returning to interoperable metadata, and to the case for treating negative results as shared infrastructure. Once again, the gap is collaboration.


Then there is economics. Capital intensity shows everywhere: Radical taking 45,000 square feet at the Brooklyn Navy Yard, Chemify's £22m of Scottish Enterprise and UK government grants for a second-generation facility in Glasgow, Lila's “AI Science Factories”. These are in effect proper lab spaces, with added investment likely going into AI-compatible equipment and autonomous pipelines. Swarms scale with compute; labs scale with square footage. Bridging the physical will take a great deal more capex.


Near Horizon: So, Where Is the Value?


My read is that value will accrue to whoever owns verification and throughput, not the labs with the strongest models.


The execution half of the scientific method — form a hypothesis, test it, discard it, go again — is being commoditised in real time. Any frontier model can be harnessed to run that loop inside a defined space, and every improvement in leading models will translate to better core intelligence capabilities. The parts that remain scarce are really everything around the loop: instruments agents can interface with, measurement fast enough to keep pace, metadata good enough for a result to be replicable across sires, and a scientific community able to check, credit and build on what this process produces.


Orchestrated swarms of agents on formal problems are already here; the bottleneck on this angle is raw compute, and perhaps the proper coordination mechanisms. Co-scientists over literature and data are here too, and weakest precisely at synthesis. Workcells optimising known processes are on the way. The full hypothesis-to-experiment loop linked to physical platforms is years away, and yet this is what the valuations are pricing in the most.


It is also where the split in capital matters. Most of the hard gaps above — shared standards, interoperable metadata, negative results as infrastructure, norms for credit — are commons problems, and the commons is what public money is buying. Private labs are building proprietary loops on infrastructure they do not own and cannot build alone. The ones that earn their valuations will be those that settle on a business model — selling discoveries, licensing the loop, or running the lab as a service — before the loop itself becomes the commodity.


What we know for certain is that this is already one of the most exciting spaces to watch, and the potential ripple effects from accelerated discovery are substantial. I suspect we are still in the early phase of commercialisation models being settled and the marked parcelling itself out into distinct niches. How this shakes out is one to watch.

 
 
LVCN - London VC Network
  • LinkedIn
  • Instagram
  • Twitter
  • Youtube
  • TikTok

The information provided on this website is for general informational purposes only. It should not be construed as professional advice or a recommendation for any particular investment. We do not guarantee the accuracy or completeness of the information provided and are not liable for any losses or damages arising from the use of this website or its contents. All investment decisions should be made at your own discretion and after thorough research. We do not endorse any specific investment opportunities or companies mentioned on this website.

CPD Member Logo.png

©2026 London Venture Capital Network Ltd.

bottom of page