AI4S Timeline

A structured record of how AI is entering scientific work — from models and benchmarks to experiments, infrastructure, corrections and deployment.

Loop
Evidence
Domain
Type

2026

Europe frames a shared 2040 strategy for materials AI

A roadmap spanning 26 organizations in 14 countries aims to connect materials research, development, scale-up and industrial deployment into one continuous system.

Its significance is not another model, but a public attempt to build the coordination and shared infrastructure that private science factories cannot easily supply alone.

Vision 2040Read More · AI4S Brief →

AlphaGenome Atlas precomputes nine billion DNA variant effects

A roughly petabyte-scale public resource turns repeated genome-wide model inference into searchable infrastructure.

The durable contribution may be the public predictive layer itself: precomputation turns an expensive model into infrastructure others can query repeatedly.

Google DeepMindRead More · AI4S Brief →

GPT-5.6 Sol runs experiments on a superconducting quantum chip

An MIT researcher connects the model through Codex to laboratory software, letting agents choose parameters, operate qubit measurements, analyze results and refine the next measurement.

AI moves from writing experimental code into a live measurement loop, while noisy or ambiguous signals still reveal where expert judgment is needed.

OpenAI

OpenAI reports an AI-generated Navier–Stokes solution amid a provenance dispute

OpenAI says an internal model produced an analytical proof and Lean formalization for the Millennium Prize formulation by constructing a finite-time singularity under smooth forcing.

If the result survives independent scrutiny, the target has moved from Olympiad problems and isolated conjectures to a Millennium Prize problem; provenance and verification are now part of the capability claim itself.

OpenAI

OpenAI says coding agents have reached its automated research-intern milestone

OpenAI reports that agents can complete well-defined tasks taking skilled researchers several days, while its research organization used 3.1 agent-workdays for every human workday by mid-August.

The evidence is observed use inside a live frontier-research organization rather than another benchmark, although the productivity data remain self-reported.

OpenAI

Claude completes a computer-checked formalization of Fermat’s Last Theorem

Anthropic reports that Claude worked largely autonomously for 11 days to produce an end-to-end Lean formalization, writing 13 million lines and proving 29,500 intermediate theorems used in the final proof.

This is not new mathematics; the advance is making a vast accepted proof machine-checkable and potentially lowering the cost of verifying future AI-generated mathematics.

Anthropic

Anthropic expands discounted Claude access for scientists

A discounted usage program expands beyond biology to fields including mathematics.

Access to frontier models is becoming research infrastructure, raising the same allocation questions as compute, data and instruments.

Anthropic

Anthropic previews a common interface for agents and lab hardware

The Model Hardware Standard proposes a shared layer for instrument control, state, measurements and errors.

A shared interface could turn one-off instrument integrations into reusable infrastructure, a prerequisite for scaling autonomous labs.

AnthropicRead More · AI4S Brief →

Co-Scientist crosses from hypotheses into physical experiments

An execution-grounded extension connects Gemini agents to real experiments, including CVD protocols for two-dimensional materials and wet-lab validation of biological predictions.

The important shift is from proposing hypotheses in silico to adapting research plans to laboratory constraints, execution and empirical feedback.

arXiv

Google moves its AI Responsibility team out of DeepMind

The roughly 90-person unit is moving into Google Global Affairs at the start of September, covering frontier-model risks, CBRN evaluations, user behavior and the psychological effects of chatbots.

This is not a layoff but a governance experiment: will moving frontier-risk research closer to policy strengthen coordination, or weaken its independence and visibility into the newest models? Google says the mission, staffing, compute and DeepMind access will remain, while model behavior, privacy and security teams stay inside the lab.

WSJ

Benchmark finds agents complete only a minority of full science workflows

Across 97 tasks in six domains, strong partial progress often fails to translate into completed research workflows.

The result warns against rewarding partial fluency: useful research agents must carry a workflow across the finish line.

arXiv

Anthropic reports Claude-designed protein binders working in the lab

External wet-lab partners validate binders for 14 of 15 targets in a model-assisted design workflow.

External wet-lab results make this more than a model demo and test whether a general system can compete inside a specialized design workflow.

Anthropic

ASI-Bench tests agents without step-by-step research guidance

Average performance drops sharply when agents must choose their own research methods rather than follow prescribed steps.

It isolates a central weakness of AI scientists: executing a method is easier than choosing the right method and knowing when to change course.

arXiv

Inherent introduces an agent trained to reproduce research papers

The startup evaluates its model on Replica, a benchmark of end-to-end paper reproduction tasks.

Replication is a useful intermediate target: harder than summarizing a paper, but more checkable than claiming open-ended discovery.

arXiv

Claude improves a longstanding bound related to the Riemann hypothesis

An unreleased Claude research model raises the known lower bound for the fraction of Riemann-zeta zeros on the critical line from 41.6% to 67.2%, with a formally checkable proof.

Unlike Olympiad performance, this is claimed novel mathematics produced while exploring an unsolved problem, without claiming to solve the Riemann hypothesis itself.

Anthropic

WeatherNext Cyclones gains roughly a day of forecast lead time

A Nature paper reports state-of-the-art forecasts for cyclone track, intensity and wind structure, with at least a day of average lead-time advantage and ensembles of up to 1,000 members.

This is unusually strong AI4S evidence because forecasts are compared with mature operational systems and then continuously checked against the physical world.

Nature

Jeff Dean leaves Google to found Discovery Loop

Dean, Sanjay Ghemawat, Quoc Le and Oriol Vinyals form a public-benefit company to automate scientific and engineering loops.

The personnel move concentrates rare systems, model and science expertise in an independent company built around closing discovery loops.

Discovery Loop

Demis Hassabis moves from DeepMind CEO to Alphabet Chief Scientist

Koray Kavukcuoglu takes operational leadership of Google DeepMind while Hassabis shifts toward long-range scientific direction.

The change matters because the lab behind many AI4S landmarks is separating long-range scientific direction from operating control.

Google

FrontierMath adds 50 genuinely unsolved research problems

Epoch AI expands the benchmark beyond private-answer questions to problems with no known solution at release.

It changes what evaluation asks: not whether a model can uncover a hidden answer, but whether it can advance a problem with no answer yet.

Epoch AI

Flex-Cat closes a catalyst-optimization loop and scales the result tenfold

The autonomous laboratory runs 680 experiments across three catalyst-optimization campaigns and transfers selected results from 2 mL discovery runs to a 20 mL reactor.

The important step is not autonomous search alone, but translation from miniaturized discovery experiments toward process-relevant conditions.

Nature Communications

AlphaProof Nexus solves open Erdős problems

A language-model and Lean search system reports solutions to nine open problems with machine-checked proofs.

Machine-checked proofs make this a meaningful step from recovering known solutions toward producing genuinely new mathematics.

arXiv

Robin links hypotheses, experiments and updated hypotheses

The system combines literature search, hypothesis generation, experimental strategy and data analysis; human researchers run the physical experiments and return the data.

Unlike one-shot hypothesis agents, experimental evidence returns to the system and changes what it proposes next.

FutureHouse

AI co-scientist reaches peer-reviewed publication

A Nature paper reports multiple biomedical case studies and experimental follow-up for hypotheses produced by the multi-agent system.

Peer review and documented follow-up make the claim more inspectable than the original product preview, even if broad autonomy remains unproven.

Nature

Evo 2 brings million-base genomic context into an open foundation model

Published in Nature, Evo 2 is trained on roughly nine trillion DNA base pairs across all domains of life with a one-million-token context window.

Genome modeling moves beyond local sequence effects toward chromosome-scale context, while reusable models and resources turn the paper into scientific infrastructure.

Nature

Gemini Deep Think extends Olympiad performance across sciences

Google reports gold-medal-level results in physics and chemistry Olympiads alongside stronger general reasoning scores.

The breadth matters, but so does the boundary: Olympiad success is evidence of reasoning, not yet evidence of autonomous research.

Google DeepMind

Shanghai AI Laboratory open-sources Intern-S1-Pro

The trillion-parameter mixture-of-experts scientific model activates 22 billion parameters per query and targets scientific reasoning.

It makes scientific foundation models an open ecosystem contest, allowing capability claims to be inspected beyond a single lab.

Shanghai AI Laboratory

OpenAI launches Prism for scientific writing

The LaTeX workspace combines drafting, citations and model assistance and is made free to ChatGPT users.

The consequential move is into the daily research surface, where assistance can shape how evidence is written, cited and shared.

OpenAI

2025

FrontierScience exposes the gap between exam and research performance

OpenAI reports strong results on Olympiad-style questions but much lower scores on open-ended research tasks.

The useful result is the gap itself: solving bounded problems does not yet transfer cleanly to open-ended research.

OpenAI

The United States launches the Genesis Mission

An executive order directs the Department of Energy to connect federal scientific data, compute and facilities around AI-driven research challenges.

AI4S becomes a national infrastructure and competitiveness program rather than only a lab or company strategy.

White House

Edison Scientific spins out of FutureHouse

A for-profit company is created to commercialize scientific agents while FutureHouse remains a nonprofit research lab.

The field experiments with separating public-interest research from commercial deployment.

FutureHouse

Lila Sciences raises a $350 million Series A

The round brings total announced funding to $550 million to scale AI Science Factories and external partnerships.

Capital formation accelerates around closed-loop experimentation as a new compute-and-data stack.

Lila Sciences

Periodic Labs launches with a $300 million seed round

Former OpenAI and DeepMind researchers form a company focused on AI-driven physical experimentation and materials discovery.

A record-sized seed round signals investor belief that experimental infrastructure can be a standalone AI platform.

Periodic Labs

AI systems reach gold-medal standard at the Mathematics Olympiad

OpenAI and Google DeepMind report natural-language solutions reaching the IMO gold threshold.

Mathematical reasoning advances beyond specialized geometry and formal-only pipelines into broader problem solving.

Google DeepMind

AlphaGenome predicts regulatory effects across long DNA sequences

DeepMind’s model predicts regulatory activity and variant effects directly from DNA sequence, pushing foundation modeling beyond structure toward noncoding function.

The harder question is no longer only whether sequence can be modeled at scale, but whether those predictions remain reliable across biological contexts and prospective use.

Google DeepMind

Google launches Weather Lab with an experimental cyclone model

The system produces 50 possible cyclone trajectories up to about two weeks ahead with input from operational forecasters.

AI forecasting moves from retrospective papers toward a live interface shaped by real forecasting institutions.

Google DeepMind

An AI-discovered drug reaches randomized Phase 2a evidence

A randomized, double-blind Phase 2a trial reports safety and efficacy results for rentosertib, a generative-AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis.

Once an AI-discovered molecule reaches controlled human trials, clinical safety and efficacy become more important evidence than the sophistication of the discovery model.

Nature Medicine

AlphaEvolve generalizes the generator–evaluator discovery loop

Gemini proposes programs while automated evaluators select and evolve them across algorithms, chip design and data-center scheduling.

A reusable discovery pattern emerges wherever outcomes can be measured cheaply and automatically.

Google DeepMind

FutureHouse opens scientific agents to the public

Specialized agents for literature search and scientific reasoning move from internal research into a public platform.

Scientific-agent capability becomes a usable service rather than only a paper result.

FutureHouse

Isomorphic Labs raises $600 million

Its first external financing is directed toward internal oncology and immunology programs and partnered drug discovery.

The AlphaFold-to-drug thesis attracts large-scale outside capital before clinical proof.

Isomorphic Labs

Lila Sciences emerges with $200 million for AI science factories

Flagship Pioneering unveils a company combining scientific models, autonomous labs and new experimental data generation.

Capital begins funding the experimental loop itself, not only software models or drug pipelines.

Flagship Pioneering

MatterGen generates crystals conditioned on target properties

Microsoft's diffusion model proposes stable inorganic materials and includes physical synthesis of a generated candidate.

Materials AI shifts from screening a fixed list to inverse design: asking what material should exist.

Nature

2024

GenCast turns weather prediction into probabilistic ensembles

The diffusion model produces ensembles up to 15 days ahead and outperforms ECMWF ENS on 97.2% of evaluated targets.

The field moves from one best forecast toward distributions of possible futures useful for decisions under uncertainty.

Google DeepMind

Recursion completes its combination with Exscientia

Exscientia becomes a wholly owned Recursion subsidiary, combining automated biology, chemistry and clinical pipelines.

AI-native drug discovery enters a consolidation phase and tests whether scale can improve translation economics.

Recursion

AlphaQubit improves quantum-error decoding

A transformer-based decoder reduces errors on Google's Sycamore experiments but remains too slow for real-time use.

A strong result arrives with a concrete systems bottleneck, illustrating the distance from benchmark gain to deployment.

Nature

The Chemistry Nobel recognizes protein design and AlphaFold

David Baker receives half the prize for computational protein design; Demis Hassabis and John Jumper share the other half for structure prediction.

AI-enabled prediction and design receive the strongest institutional recognition in science.

Nobel Prize

The Physics Nobel recognizes neural-network foundations

John Hopfield and Geoffrey Hinton receive the prize for foundational discoveries enabling machine learning with neural networks.

The scientific establishment places machine learning's foundations inside the history of physics.

Nobel Prize

PaperQA2 reports superhuman scientific literature retrieval

FutureHouse's agent exceeds PhD and postdoctoral biologists on selected literature-search tasks and releases its code.

Scientific agents begin by attacking the information bottleneck, with citation-aware retrieval as a measurable task.

FutureHouse

AlphaProteo designs experimentally successful protein binders

DeepMind reports successful binders across seven targets, including VEGF-A, alongside an unsuccessful eighth target.

Protein generation is judged by wet-lab hit rate and affinity rather than structural plausibility alone.

Google DeepMind

Recursion and Exscientia agree to combine

Two of the most visible public AI-drug-discovery companies agree an all-stock transaction.

The sector begins consolidating as platform promises meet the capital intensity and timelines of clinical development.

Recursion

AlphaProof and AlphaGeometry 2 reach IMO silver-medal standard

The systems score 28 of 42 points, one point short of gold, though some solutions take days rather than competition hours.

Formal verification provides cheap, strict feedback for large-scale mathematical search.

Google DeepMind

AlphaFold 3 predicts interactions across biomolecules

The model extends beyond proteins to complexes containing DNA, RNA, ligands and modified residues.

Structure prediction moves closer to the interaction problems central to drug discovery, while restricted access triggers an openness dispute.

Nature

AlphaGeometry approaches Olympiad gold-medal performance

The neuro-symbolic system solves 25 of 30 historical Olympiad geometry problems within competition time limits.

Language-model generation and symbolic deduction form a verifiable search loop for mathematics.

Nature

Isomorphic Labs signs Lilly and Novartis collaborations

Two multi-target drug-discovery collaborations carry nearly $3 billion in potential value, excluding royalties.

Large pharmaceutical companies place substantial milestone-based bets on AI-first drug design.

Isomorphic Labs

2023

Coscientist plans and executes chemistry experiments

A GPT-4-based multi-agent system searches documentation, plans reactions, writes control code and operates a cloud laboratory.

The LLM becomes an orchestrator for a real experimental workflow rather than only a scientific text interface.

Nature

FunSearch pairs language-model generation with an executable evaluator

The system evolves LLM-generated programs against an automated evaluator and reports new cap-set constructions alongside improved bin-packing heuristics.

A reusable discovery pattern appears: models can search broadly when incorrect ideas are rejected cheaply and objectively by execution.

Google DeepMind

GNoME expands the computed landscape of stable materials

DeepMind reports 2.2 million candidate crystal structures, including 380,000 added to a set judged most stable.

Materials screening reaches industrial scale while sharpening debate over what should count as a discovered material.

Nature

A-Lab closes a materials synthesis loop

The autonomous lab runs 353 experiments in 17 days and realizes 36 of 57 target inorganic materials.

Prediction, literature-derived recipes, robotics, measurement and active learning operate as one physical loop.

Nature

GraphCast beats a leading physics-based weather system

The graph neural network outperforms ECMWF HRES on more than 90% of 1,380 verification targets and produces a ten-day forecast in under a minute.

Learned forecasting proves competitive under a mature, continuously reality-checked evaluation regime.

Science

FutureHouse launches to build an AI scientist

The Schmidt-backed nonprofit begins building scientific agents, initially focused on biology and scientific literature.

AI-for-science agent development becomes the mission of a dedicated research organisation rather than a side project.

FutureHouse

AlphaMissense maps 71 million possible human missense variants

The model classifies most possible missense variants as likely benign or pathogenic, vastly expanding the small experimentally interpreted set.

AlphaFold-derived representations move from structure prediction into genome interpretation and disease research.

Science

RFdiffusion turns protein structure prediction into protein design

A diffusion model generates novel protein backbones, with hundreds of designs tested experimentally.

The central question shifts from “what shape does this sequence have?” to “what sequence and structure should exist?”

Nature

Pangu-Weather challenges conventional global forecasting

Huawei's 3D neural network reports stronger deterministic forecasts than ECMWF's operational system while running more than 10,000 times faster.

A mature scientific field with strict operational evaluation becomes one of AI4S's clearest tests outside biology.

Nature

AlphaDev's sorting algorithms enter the C++ library

Reinforcement learning finds faster low-level sorting routines that are added to LLVM's libc++ implementation.

An AI-discovered algorithm moves from a paper into infrastructure used by developers at scale.

Nature

Google merges Brain and DeepMind into Google DeepMind

Demis Hassabis takes charge of the combined unit; Jeff Dean becomes Google's Chief Scientist.

Talent, compute and model development are concentrated inside one organisation that also houses major AI4S programs.

Google

2022

Meta withdraws the Galactica scientific language-model demo

The system produces authoritative-looking false claims and citations, and the public demo disappears after three days.

Scientific fluency is shown to be different from scientific reliability.

Galactica

ESMFold maps more than 600 million metagenomic proteins

Meta uses a protein language model to predict structures without the multiple-sequence alignments required by AlphaFold.

Protein language modeling offers a faster, complementary route to structure at unprecedented scale.

Science

AlphaTensor discovers new matrix-multiplication algorithms

Reinforcement learning finds algorithms that improve on long-standing human-designed solutions for selected matrix sizes.

Scientific search becomes a game with an executable evaluator, a pattern later generalized by AlphaDev and AlphaEvolve.

Nature

AlphaFold DB expands beyond 200 million structures

The database grows from around one million structures to predictions covering nearly every catalogued protein.

Prediction at planetary scale changes protein structure from a scarce result into a queryable layer of biology.

Google DeepMind

A reinforcement-learning system controls plasma in a tokamak

A single neural controller commands the magnetic coils of Switzerland's TCV reactor and sustains difficult plasma shapes.

AI moves beyond offline prediction and directly operates a real scientific instrument.

Nature

2021

Alphabet launches Isomorphic Labs

A new commercial company led by Demis Hassabis is created to rebuild drug discovery around an AI-first approach.

AlphaFold's scientific success produces a dedicated company aimed at therapeutic translation.

Isomorphic Labs

The AlphaFold Protein Structure Database opens

The first release makes roughly 350,000 predicted structures searchable, including the human proteome.

Model inference becomes a public research resource that can be reused without rerunning the model.

EMBL-EBI

AlphaFold 2 is published and released as open source

The architecture, code and methods behind the CASP14 result become available to the research community.

The breakthrough starts turning into shared scientific infrastructure rather than a single demonstration.

Nature

2020

AlphaFold 2 changes the protein-structure problem at CASP14

DeepMind reports accuracy competitive with experimental structures in the blind CASP assessment.

A decades-old scientific bottleneck becomes a tractable learned-inference problem.

Google DeepMind