Frontier question · Choose
Does the AI Scientist win through intelligence — or a new economics of experimentation?
AI4S.fyi synthesis · Current answer
AI Scientist gains may come from better experiment selection, faster execution and lower cost at the same time.
Current cases often report these changes together, making their individual contributions hard to attribute. Comparisons should examine results under equal budgets and the total time and cost required to reach the same objective.
Confidence: High that throughput has improved. Low on attribution, at every one of the four sources below. This is not a question where the evidence is mixed — it is a question where the evidence is confounded, which is a different and more fixable problem.
Why it matters
If the advantage is intelligence, it should transfer across domains and compound: each campaign makes the next one cheaper everywhere. If it is economics, the advantage is just as real but buys something different — it is reproducible by anyone willing to spend, it has to be re-earned in every new vertical, and it does not accumulate into a model that gets smarter. The two stories imply opposite capital strategies and opposite moats.
Four sources of advantage, and which side each belongs to · 7 entries
Decide
Selection quality
Intelligence · 1 key
Throughput
Economics · 1 key
Cost per result
Economics · 2 key
Regions humans avoid
Intelligence · 2 key
Interact
This decomposition is by confound, not by method: these are the four things you would have to hold fixed, one at a time, to tell the two explanations apart. No published result holds any of them fixed.
Source one · attributed to intelligence
Selection quality at a fixed budget
Does the model choose better experiments than a human would, given the same number of shots?
AssessmentOne prospective test, in one alloy family. This is the claim the intelligence story depends on, and it is the least tested of the four.
Confidence: low · 1 key entry
Key evidence
HEA05
→ IntelligenceProspective · Physical · Peer reviewed
2026.08
Bearing on this question
Four adaptive rounds prospectively tested eight new candidates — the only entry on this page where selection was tested forward rather than claimed retrospectively.
Does not establish
No control arm selecting by other means, no cross-family or cross-lab replication, and no fixed-budget comparison against experienced human selection.
Open evidence →
Further evidence
1 entry · anecdotal
2026.08
Radical AI interview — human scientists occasionally submit competing compositions and are rejected by the model; told as a story, not a measurement Testimony
Intelligence
Source two · attributed to economics
More useful experiments in the same time?
Whether more shots per unit time explains the result on its own.
AssessmentRoughly 1,200 samples in five to six months versus about 500 in twelve months implies a crude 4.8–5.8× rate ratio, not a demonstrated tenfold gain. The projects may differ in scope and depth, so this is not a controlled efficiency comparison.
Confidence: moderate-high on the throughput itself; it provides no support to the intelligence side · 1 key entry
Key evidence
Radical AI — alloy campaign throughput
→ EconomicsCompany self-report · Industry comparator
2026.08
Bearing on this question
Roughly 1,200 alloys made and characterised in five to six months, against an industry comparator — a DARPA and GE Aerospace programme — of about 500 alloys in twelve months. Current throughput is 8–20 experiments a day with a stated target of 100.
Why this is the crux
A tenfold increase in shots will produce more discoveries even with unchanged selection quality. Any claim that the model is smarter has to survive subtraction of this effect, and no published result subtracts it.
Does not establish
Whether the comparator programme's experiments were of comparable depth; whether hit rate holds as throughput scales.
Open interview →
Source three · attributed to economics
Lower total cost for the same objective?
Not cost per experiment — cost per experiment that changed something.
AssessmentPer-experiment costs are published; cost per useful discovery is not, by anyone. Rising throughput does not imply falling unit cost of a result: if hit rate declines as volume grows, it can rise.
Confidence: low · 2 key entries, neither externally verified
Key evidence
Radical AI — per-experiment cost
→ EconomicsOn-the-record estimate
2026.08
Bearing on this question
Roughly $60–300 per alloy experiment, depending on the elements used. A rare public number for the denominator of the economics argument.
Does not establish
A conversational estimate, not audited. Says nothing about how the figure moves as the lab scales toward its stated target.
Open interview →
Lila Sciences — the zero-FTE startup claim
→ EconomicsCompany self-report
2026.07
The claim
Two to three people with the model and platform can do five years of biotech work in six months for about a tenth of the investment.
Bearing on this question
The most explicit statement anywhere that the advantage is economic rather than cognitive — made by a company whose stated core asset is nonetheless the model.
Does not establish
No external verification, no disclosure of how the comparison baseline was constructed.
Open interview →
Source four · attributed to intelligence
Valuable new regions explored?
Not better choices within the usual space — choices outside it.
AssessmentThe most interesting evidence the intelligence side has, and it is entirely internal. Two independent companies report the same shape of result, which is worth something; neither ran a fixed-budget comparison, which limits what.
Confidence: moderate on the phenomenon, low on the attribution · 2 key entries
Key evidence
Radical AI — the elemental-families overlay
→ IntelligenceCompany self-report · Internal chart
2026.08
Bearing on this question
Overlaying the elemental combinations explored in published literature with those explored by the system shows the model entering families never published on. Asked why they had avoided them, the company's own scientists gave reasons — it would evaporate, it would not cast, the microstructure would not hold — that turned out to be wrong. This is a specific, checkable mechanism for an advantage that is not throughput.
Does not establish
The chart is internal and unpublished. Avoiding a region is not the same as being wrong to avoid it in expectation; the successes are reported, the failures in those families are not.
Open interview →
Lila Sciences — the non-platinum-group electrocatalyst
→ IntelligenceCompany self-report
2026.07
Bearing on this question
A specialist with roughly forty papers on the topic judged the model's suggestions first boring, then wrong; they became the company's best-performing non-platinum-group catalysts. Independently reported, it matches the Radical result closely enough to suggest a real phenomenon rather than one company's story.
Does not establish
Single case, company-reported, no denominator — how many equally "wrong-looking" suggestions failed is not disclosed.
Open interview →
A methodological note. The defect here is the same one that prevents the information-fidelity question from being settled: MIST used a repaired tokenizer at 1.8B parameters but changed model scale, training data and compute budget at the same time, so the tokenizer's contribution cannot be isolated. Here, every reported win changes selection, throughput and cost together, so intelligence cannot be isolated from economics. Both questions are blocked by the same missing practice — holding one term fixed.
Counter-evidence and cross-pressure
The case that it is economics all the way down
Two of the three entries below come from people running these systems, and both point away from the intelligence explanation. The third points the other way and is the weakest evidenced.
Joseph Krause — experiment-constrained, not compute-constrained
→ EconomicsExpert testimony
2026.08
The claim
"We're not compute constrained in the materials industry. We're experiment constrained."
Why it matters here
A direct statement from an operator that the binding constraint is the price and pace of physical information, not the quality of reasoning about it.
Open interview →
Max Welling — automation is vertical-specific
→ EconomicsExpert testimony
2026.02
The claim
Completely automating one problem does not carry over to the next: the experimental setup differs, the characterisation instruments differ, and the platform's models have to be retrained and fine-tuned.
Why it matters here
If this holds, there is no accumulating cross-domain intelligence to win with — only per-vertical economics, re-earned each time. It is the sharpest available argument against the intelligence story, and it comes from someone whose company would benefit from the opposite being true.
Open interview →
Lila Sciences — "breadth gives us depth"
→ IntelligenceCompany self-report
2026.07
The claim
Training across more scientific domains reduces the domain-specific data needed for any one of them, sometimes to near zero — the direct contradiction of the entry above.
Does not establish
No public cross-domain controlled benchmark. Supported internally by cases such as reusing small-molecule chemistry reasoning on metal-organic frameworks.
Open interview →
AI4S.fyi synthesis · Current gap
Mechanism-level feasibility is far better supported than system-level attribution. The missing study is cheap by the standards of this field: the same laboratory, the same time window, the same number of experiments, selected two ways.
What would change our view
- SelectionA fixed-budget head-to-head: N experiments chosen by the model against N chosen by experienced humans, same lab, same window, outcomes scored blind.
- CostA published curve of cost per useful discovery as throughput scales, showing whether hit rate holds.
- CoverageThe elemental-families overlay, or an equivalent, reproduced by a group that is not the company reporting it — including the failures in those regions.
- CounterAny demonstration that a model trained in one vertical transfers to another without retraining, which would settle the Welling–Lila disagreement directly.
Related reading
Radical AI on throughput, cost per experiment, and why the binding constraint is the price of information.
CuspAI's Max Welling on why automation does not transfer between verticals.
Blocked by the same missing practice: holding one term fixed while the others move.