AI4S.fyi Evidence Framework
How we think about AI for Science
The hardest question is not whether AI can make a prediction. It is whether a scientific system can learn from the world.
THE LOOP
Scientific learning requires four actions
Understand what we know
This contains two distinct abilities. Reading existing science recovers what papers, figures and databases already say. Modeling the system predicts states that have not yet been observed.
Choose what to try
Candidates usually exceed the experimental budget. A useful decision system must do more than rank answers: it must justify where the next dollar, hour or experiment should go.
Test against reality
Simulation and proxies can filter candidates, but new evidence often appears only after synthesis, execution and measurement. Measurement itself can introduce noise, drift and failure.
Learn from the result
A loop closes only when a new observation changes the model or next decision. The stronger test is whether that change makes the next round better.
STRESS TESTS
Four ways a promising result can fail
It works—until the setting changes.
GeneralizationDoes the result survive a new material, molecule, laboratory or data distribution?
It looks convincing—but the evidence may be weak.
TrustDo predictions correspond to reality, and can measurements be calibrated, inspected and repeated?
It works once—but not at useful scale.
ScaleDo throughput, failure rate, time and unit cost support sustained operation?
Every component works—but the system does not.
IntegrationCan models, software, instruments, automation and human exceptions form a stable workflow?
EVIDENCE
Not all evidence touches reality in the same way
- 01
Retrospective or simulated
Uses existing data or simulated environments. Useful for comparison, but the dataset can be mistaken for reality.
- 02
Prospective in silico
Makes a prediction before the answer is known. This limits hindsight, but may still rely only on computational proxies.
- 03
Physical validation
Real synthesis, measurement or clinical observation places the claim in contact with the physical world.
- 04
Closed-loop experimentation
Results return to the system and change the next model or experimental choice.
- 05
Field or production use
The system enters real scientific, clinical or industrial workflows and faces long-term operation, recovery and cost.
Moving closer to reality does not automatically make a claim correct. It exposes the system to different—and usually harder—ways to fail.
AGENCY
Who actually did the work?
“AI discovered” often compresses several roles into one phrase. Separating them shows how far autonomy actually travelled.
People define the problem, constraints and success criteria, and adjudicate ambiguous results.
Models propose hypotheses, predict outcomes or select the next action.
Robots and software execute a predefined workflow.
People, models and automation jointly make decisions and carry them out.
CLAIMS
Scientific claims change
We do not label an entire paper “true” or “false.” Important claims are tracked separately and updated as new evidence appears.
CANON
Canon does not mean “best paper”
A work can become Canon because the field cannot explain its current methods, validation standards, failures or controversies without it. Canon is not an honor roll; a canonical claim may later be corrected or even refuted.
THE LIBRARY
How the library evolves
Worth watching.
Actively changing the field.
Impossible to route around.
No longer active, but historically preserved.
CONTEXT
Science does not happen in a vacuum
Funding, institutions, policy and infrastructure shape which questions get pursued and which systems can be built. We track that context while reading it separately from evidence that a scientific capability works.
What we are ultimately trying to see
Can AI move from understanding what is known, to choosing what to test, to learning from reality—and do so reliably enough to change how science is done?
Explore the evidence →