Frontier question · Model systems

Is information lost during measurement, extraction, or encoding a hard ceiling on scientific AI, or can scale buy it back?

AI4S.fyi synthesis · Current answer

Larger models can mitigate some data shortages, but do not guarantee recovery of information absent from the input.

The remedy depends on whether information was not measured, incompletely extracted, merged during encoding or distorted in the reference labels.

Updated 2026-09-10. Model scale, training-data scale and inference coverage are assessed separately; entry counts do not substitute for evidence strength.
Why it matters

Information recovery and predictive improvement are different. If two objects are encoded identically, a model cannot reliably distinguish them from that input alone, although context and correlated features may still improve a probabilistic prediction. Experimental and computational labels may follow different paths, so the four perspectives below are not a mandatory linear pipeline.

The INTERACT → MODEL arc, magnified · 10 entries
Interact
Measurement
2 load-bearing
1 further
Documentation
2 load-bearing
0 further
Encoding
2 load-bearing
1 further
Training labels
2 load-bearing
0 further
Model

Counts are not weights. These four gates magnify one arc of the site's learning loop: what happens between the world being measured and a model beginning to learn.

Gate one of four · toward MODEL

Measurement

What was never observed in the first place.
AssessmentScale extends reach, not fidelity. More computation covers more of the space, but it propagates what the base model could not see rather than correcting it.
证据边界:moderate · 2 load-bearing entries
Key evidence
AlphaGenome Atlas
→ RefinesPeer-reviewed base model
2026.09
What it does
Precomputes genome-wide variant effects into a queryable predictive map.
Bearing on this question
Nine billion precomputations extend coverage across the map but carry the base model's blind spots with them—a direct separation of scale of inference from scale of fidelity.
Does not establish
Most records are predictions; no system-level independent replication or clinical validation.
Open evidence →
CuspAI — the laboratory as a processor
→ CeilingExpert testimony
2026.02
Position
Max Welling proposes adding a new information source rather than scaling the existing one: treat the lab as a physics processing unit and let nature perform the computation. His stated bottleneck is that simulation is not predictive enough to substitute for measurement.
Bearing on this question
The strongest available restatement of the ceiling position: if more scale were the answer, a materials company would not need to build a laboratory at all.
Does not establish
Testimony, not demonstration. The platform's throughput and hit rates are undisclosed.
Open interview →
Further relevant evidence
1
  • 2026.08
    Radical AI interview — “we're not compute constrained, we're experiment constrained”
    Testimony
    Ceiling
Gate two of four · toward MODEL

Documentation & extraction

Recorded once, but locked inside papers, figures and scans.
AssessmentRecoverable in principle—the one gate where effort does buy the information back. The information still exists; getting it out is an engineering problem, and it is not solved yet.
证据边界:high on reversibility, low on completeness · 2 load-bearing entries
Key evidence
MinerU.Chem
→ Scale / effortPublic system
2026.06
What it does
Recovers chemistry objects and reaction relationships from papers into traceable data.
Bearing on this question
Loss at this gate is reversible, which is exactly why it is the wrong gate to generalize from. Optimism about the whole question usually comes from looking only here.
Does not establish
No document-level, database-ready accuracy without human correction.
Open evidence →
OCSR / MolRecBench
→ RefinesPublic benchmark
2026.06
Bearing on this question
Recovery is real but far from free: abbreviations, drawing conventions and complex structures still defeat it. Effort buys information back here, but neither cheaply nor completely.
Does not establish
The sample does not cover all eras, patents, scan qualities or chemistry subfields.
Open evidence →
Gate three of four · toward MODEL

Encoding

What survives the trip into the model's input format.
AssessmentHard ceiling until the encoder is repaired. Scale cannot help, because the information never enters the model. Every fix on record was a change of representation, not an increase in scale.
证据边界:high · unchanged since this question opened · 2 load-bearing entries
Key evidence
Smirk
→ CeilingPeer reviewed · JCIM
2026.01
What it does
Preserves rare OpenSMILES combinations that closed vocabularies collapse into a placeholder.
Bearing on this question
The cleanest case on this page. The loss is silent, training does not error, and no quantity of additional molecules recovers a chirality tag that was already replaced. Repairing the tokenizer did.
Does not establish
No universal downstream gain; it does not restore physical information absent from SMILES. MIST used it at 1.8B parameters but changed scale, data and budget together, so it does not isolate the tokenizer's contribution.
Open evidence →
Radical AI — “Composition is not a material”
→ CeilingExpert testimony
2026.08
Bearing on this question
A composition string cannot encode synthesis route, microstructure or processing—the things that decide whether an alloy performs. This is the encoding gate in a field that has no equivalent of SMILES at all.
Does not establish
Testimony from a company whose business depends on the claim being true.
Open interview →
Further relevant evidence
1
  • 2026.02
    CuspAI interview — materials has no canonical string representation to repair in the first place
    Testimony
    Ceiling
Gate four of four · toward MODEL

Training labels

What the teacher itself had already lost.
AssessmentHard ceiling. Scaling gives you more of the teacher, not a better teacher. Breaking it requires changing the teacher—higher-accuracy methods, or experiment.
证据边界:high on mechanism, low on coverage · added 2026.09 · the newest and thinnest gate
Key evidence
Interatomic potentials and the DFT ceiling
→ CeilingField practice
2026.09
Bearing on this question
Machine-learning interatomic potentials learn DFT-computed energies and forces, so their ceiling includes DFT's systematic failures on van der Waals interactions, strongly correlated systems, band gaps and reaction barriers. Contrast AlphaFold, which learned from experimentally determined structures: two celebrated successes on opposite foundations.
Why scale fails
Scaling produces more DFT data, not better DFT.
Does not establish
No controlled comparison against an experiment-trained equivalent exists.
Open evidence →
Lila Sciences — sim2real as the one bottleneck
→ CeilingExpert testimony
2026.07
Bearing on this question
Asked which single bottleneck he would remove, Rafa Gómez-Bombarelli chose sim2real: simulations are not predictive enough, and models trained on those data cannot close the gap because the approximations are not good enough. He then points to the asymmetry—AlphaFold was trained on experiments, and materials has no equivalent.
Does not establish
Testimony, not demonstration.
Open interview →

Alternative explanations and limitations

When scale may still improve prediction

Cross-domain transfer may improve prediction without recovering information already lost. Current company reports change data, training and tasks together, so they do not isolate model scale.

Lila Sciences — “breadth gives us depth”
→ ScaleCompany self-report
2026.07
The claim
Training across more scientific domains reduces the domain-specific data needed for any one of them, in some cases to near zero.
Does not establish
No public cross-domain controlled benchmark. Internal support includes reusing small-molecule chemistry reasoning on metal-organic frameworks.
Open interview →
Radical AI — MATRIX transfer result
→ ScaleCompany preprint
2026.08
The claim
A vision-language model trained on experimental lab images gained roughly 5–16% on some general scientific-reasoning benchmarks, with transfer into biology.
Does not establish
The result is benchmark-dependent, and the company reports no gain on mathematical reasoning.
Open evidence →
Max Welling — the inductive-bias concession
→ ScaleExpert testimony
2026.02
The claim
Ultimately it is a trade-off between data and inductive bias, and architectures should always be built to scale.
Why it matters here
This frames a testable trade-off between physical priors and data scale without claiming that scale recovers information already lost.
Bounded by
It concerns compensating for scarce data, not destroyed data—so it narrows the ceiling claim without overturning it.
Open interview →

AI4S.fyi synthesis · Current gap

The cases observe several kinds of loss separately, but do not establish every transition for the same object. An end-to-end audit is still needed, separating reference-method, learning, sampling and experimental-condition errors.

What would change the assessment?
  • All gatesAn end-to-end reversibility test for one scientific object across all four gates.
  • EncodingA controlled comparison isolating gains from model scale versus repaired input representation.
  • LabelsAn interatomic potential trained on experimental labels, benchmarked against a DFT-trained equivalent.
  • CounterA public cross-domain controlled benchmark for breadth gives depth, produced by a group that is not selling the platform.

Related reading
CuspAI's Max Welling on materials search, physical priors, and why automation starts from expert workflows.
Lila Sciences on experiments as a training signal—and sim2real as the bottleneck they would remove first.
Radical AI on why composition is not a material and why the binding constraint is the price of an experiment.