AI FOR SCIENCE · ARGUMENTS

AI4S Commentary

Arguments worth tracking across AI for Science, from first-hand accounts to technical challenges and formal scientific exchange.

View
Area
Loop
Issue
Voice

Latest commentary

A dated feed of claims, interpretations and challenges worth following now.

36 / 36

2026

Anthropic Just Gave AI Agents a Driver for Your Lab Instruments

ArgumentStandard lab interfaces make instrument control easier, but API access does not give an agent a model of physical state or safe operation.

Why readRead it for the Genentech foam failure: the software call was correct while the experiment was physically wrong. It makes the next bottleneck unusually concrete.

Scientific Creativity through Analogical Reasoning, Virtual Labs, and the Physics of Agents

ArgumentA useful AI scientist needs non-obvious analogy, productive disagreement and a model of environmental dynamics, not only fluent reasoning in language.

Why readZou shifts the question from whether agents can execute a workflow to whether they can generate the kind of disagreement from which new science begins.

Measuring the Frontier of Scientific Capability in the Age of AI

ArgumentAI scientists should be evaluated in hidden, interactive worlds where evidence can change hypotheses, not with static question-answer benchmarks.

Why readOne of the clearest proposals for what to build next: limited budgets, real feedback and traces that reveal whether an agent actually updates.

AI & CBRN Risks

ArgumentAI risk depends on where it changes a threat pathway, so CBRN assessment must map concrete steps and physical bottlenecks rather than count information alone.

Why readRead after Olvera. One says the physical barrier remains high; this one asks which specific barrier AI might lower. The disagreement is about where intervention matters.

Why read together: Both begin with the same fact: turning an idea into physical reality is hard. Read together to see how one bottleneck can limit both scientific benefit and harmful capability.

Why AI Won’t Cure Cancer Anytime Soon

ArgumentAI may compress target and molecule discovery, but toxicity, pharmacokinetics, trial design, recruitment and regulation still return evidence on year-long clocks.

Why readIt turns the promise of faster biomedicine into a sequence of named stages. That makes clear which clock a model can shorten and which clocks it cannot yet touch.

Featured exchangeAll exchanges

You Cannot Vibe Science

ArgumentPlausibility is not enough in experimental science; automation needs inspectable methods, constraints and provenance because mistakes enter the physical world.

Why readRickerby actually let an AI choose work down to reagent and well. That makes this a report from the boundary, not a forecast about it.

AI Agents Can’t Yet Do Open-Ended AI Research

ArgumentStrong benchmark scores do not yet transfer to open-ended research where problem choice, judgment and changing direction are part of the task.

Why readTheir method matters more than the headline: inspect how the number was produced, how much scaffolding contributed and what the agent chose without help.

Anecdotes Everywhere, Evidence Almost Nowhere

ArgumentClaims that AI is transforming productivity remain dominated by anecdotes, while rigorous evidence is sparse and unusually hard to interpret.

Why readAfter the withdrawn MIT preprint, this becomes an editorial standard for every spectacular productivity number: ask what was actually measured.

Conjecture Machines

ArgumentIf agents make hypotheses cheap, validation, agent-ready data and peer review become the new scientific bottlenecks.

Why readA rare piece that asks what funders and institutions should build once idea generation is no longer scarce.

Why I Left Google DeepMind

ArgumentPublic commitments to responsible AI can become ineffective when internal incentives make dissent costly and external accountability weak.

Why readIts value is specificity, not neutrality. Turner separates public evidence from private exchanges that readers cannot independently verify.

Can AI Make Scientific Breakthroughs?

ArgumentModels can solve well-specified problems without possessing the tacit knowledge, experience of failure and judgment needed to identify which problems are worth specifying.

Why readA structural challenge to today’s science benchmarks: they score answers to questions that have already survived the hardest act of scientific judgment.

AI Can Help Plan a Bioweapon. Building One Is Still Hard.

ArgumentInformation access is not the decisive bottleneck for bioweapons; tacit wet-lab knowledge and physical execution still dominate the path to harm.

Why readThe argument mirrors AI4S optimism: the same difficulty of turning ideas into physical reality limits both harm and scientific benefit.

Read with · Reading order 1/2 AI & CBRN Risks

Why read together: Both begin with the same fact: turning an idea into physical reality is hard. Read together to see how one bottleneck can limit both scientific benefit and harmful capability.

RLVR Might Be Disproportionately Bad at Science

ArgumentRLVR works in mathematics and code because verification is cheap and reliable; most important scientific questions lack that condition.

Why readThe most direct rebuttal to the claim that success in math and code should transfer to science. The verifier, not only scale, may decide the outcome.

Science Needs AI Data Stocktakes

ArgumentGovernments cannot plan AI for science without first mapping which datasets exist, who controls them and whether machines can actually use them.

Why readLess glamorous than a model launch, but far more concrete about the missing public infrastructure on which many launch claims depend.

Request for Proposals: The Launch Sequence

ArgumentTransforming ideas into scoped projects, funded teams and executable institutions is itself a piece of research infrastructure.

Why readIt treats metascience as something to build and test, making it a useful answer to the question of where the bottleneck moves when hypotheses get cheaper.

2025

Getting Started in BioML Research & Engineering

ArgumentUseful BioML work depends more on strong foundations, comfort with messy data and rapid question testing than on memorizing large amounts of biology.

Why readA practical map of the hybrid scientific and engineering talent AI4S labs need, written by someone hiring and building in the field.

2024

AI Does Not Make It Easy

ArgumentBetter structures and molecular suggestions do not remove the slow biological, pharmacokinetic, toxicological and clinical failure modes that dominate drug development.

Why readA representative dispatch from medicinal chemistry rather than a generic blog index. Lowe’s durable point is that structure can be difficult without being the rate-limiting step.

Machines of Loving Grace

ArgumentPowerful AI could compress fifty to one hundred years of biological and medical progress into five to ten years.

Why readThe most cited optimistic claim belongs here because it places the bet somewhere testable: is the bottleneck ideas, or experimental throughput, trials and regulation?

Open Letter on AlphaFold 3 Reproducibility

ArgumentA high-impact methods paper without runnable code cannot meet the reproducibility standard expected of scientific publication, especially when the result comes from a corporate lab.

Why readThe letter matters because it tested whether journal norms still bind frontier corporate research. AlphaFold 3’s later code release made the dispute consequential rather than symbolic.

Response to ‘The Perpetual Motion Machine of AI-Generated Data…’

ArgumentRapidly falling genome-sequencing costs may produce biological data at a scale that weakens the claim that data scarcity will remain fixed.

Why readRead second. It accepts much of Listgarten’s case but attacks one assumption directly. Two short pieces, one clean disagreement.

Why read together: Listgarten says science is bottlenecked by empirical data; Noble accepts that frame but challenges whether genomics will remain data-poor. Read together because the disagreement lands on a measurable future curve.

Artificial Intelligence Driving Materials Discovery?

ArgumentGNoME’s proposed structures offer scant evidence of compounds that are simultaneously novel, credible and useful without synthesis expertise and validation.

Why readThe key distinction is between computational discovery and experimental discovery. Later corrections in autonomous labs return to exactly this boundary.

Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis

ArgumentReanalysis of A-Lab’s reported products found systematic diffraction and novelty problems, so the original evidence did not establish 43 scientifically new materials.

Why readThe archive’s hardest external check: it uses the authors’ diffraction data rather than a rival forecast. The later correction adopts the same distinction between new to a platform and new to science.

The Perpetual Motion Machine of AI-Generated Data and the Distraction of ChatGPT as a ‘Scientist’

ArgumentThe fundamental bottleneck in science is data, not models, and generating synthetic data with AI cannot create the missing empirical information.

Why readRead first. It is the most compressed skeptical case on the page and explains why protein folding may be an exception rather than a template.

Why read together: Listgarten says science is bottlenecked by empirical data; Noble accepts that frame but challenges whether genomics will remain data-poor. Read together because the disagreement lands on a measurable future curve.

2023

2022

A Vision of Metascience

ArgumentScientific institutions can be deliberately redesigned and tested so the discovery ecosystem learns how to improve its own social processes.

Why readIt answers the question raised by cheap hypotheses: attention, funding and experimental capacity become allocation problems that models alone cannot solve.

2021

Protein Structure Prediction by AlphaFold2: Are Attention and Symmetries All You Need?

ArgumentAccurate structure prediction is not the same as understanding the physical and chemical principles that govern protein folding.

Why readFor readers past the popular explanation. The distinction between prediction and understanding will recur in every field that follows.

2020

AlphaFold2 @ CASP14: ‘It Feels Like One’s Child Has Left Home.’

ArgumentAlphaFold2 was not a normal increment; it effectively closed single-chain protein structure prediction as a competitive benchmark problem.

Why readStart here. Part of AlQuraishi’s own research direction was displaced, yet he was among the earliest to state the result’s full weight without hedging.

Read with · Reading order 2/2 AlphaFold @ CASP13: ‘What Just Happened?’

Why read together: The first essay captures what AlphaFold still appeared unable to solve before AlphaFold2; the second records how quickly that assessment had to be revised after CASP14.

Transparency and Reproducibility in Artificial Intelligence

ArgumentImpressive clinical AI results lose scientific value when missing code and methodological detail prevent independent scrutiny and learning.

Why readThe first full journal-standard confrontation of the AI4S era. The criticism, reply and added methods show that a dispute should end with inspectable material, not parallel statements.

Why read together: The challenge asks what scientific value remains when a result cannot be inspected; the reply and same-day addendum show whether journal pressure can turn that objection into additional material.

Reply to: Transparency and Reproducibility in Artificial Intelligence

ArgumentClinical constraints and protected data can limit full release, but a model can still be evaluated and its methods made substantially more inspectable.

Why readRead second, together with the same-day addendum. The value is not that every objection vanished, but that the exchange produced additional experimental detail.

Why read together: The challenge asks what scientific value remains when a result cannot be inspected; the reply and same-day addendum show whether journal pressure can turn that objection into additional material.

2019

The Bitter Lesson

ArgumentMethods that scale with computation eventually outperform approaches built around human knowledge and handcrafted structure.

Why readAI4S architecture debates repeatedly return here. Keep its boundary visible: the historical evidence comes from domains rich in data, while science is often data-poor.

2018

AlphaFold @ CASP13: ‘What Just Happened?’

ArgumentAlphaFold’s first CASP result was impressive but still legible as progress within the existing academic trajectory.

Why readRead before the 2020 essay. Watching one expert revise his own judgment over two years is more persuasive than any retrospective history.

Why read together: The first essay captures what AlphaFold still appeared unable to solve before AlphaFold2; the second records how quickly that assessment had to be revised after CASP14.