Frontier question · Test
Can AI move from proposing candidates to closed-loop learning in the physical world?
AI4S.fyi synthesis · Current answer
Some systems now connect candidate generation, experiments and the next selection round. The harder question is how much feedback improves decisions, and whether the gain persists and replicates.
We should examine measurement quality, objective validity, feedback updates and multi-round performance separately. A human-in-the-loop system can still learn from feedback; closed-loop learning is not the same as full autonomy.
Updated 2026-09-10. Company architecture claims, peer-reviewed experiments and independent validation are different levels of evidence; entry counts are not evidence strength.
Why it matters
A measurement is not automatically a learning signal. Noise, drift, batch effects, failed runs and convenient proxies can all teach a model the wrong lesson — and a loop that looks closed on an architecture diagram may still have a human reading a report somewhere in the middle of it. The claim "our lab is a closed loop" is four claims, and they are usually evidenced very unevenly.
The INTERACT → UPDATE → MODEL arc, magnified · 7 entries
Interact
Measured reliably
1 key
1 further
Measures the right thing
2 key
0 further
Signal returns
2 key
1 further
Next round improves
1 key
0 further
Model
Counts are not weights. These four checkpoints magnify the return path of the site's learning loop — the segment that carries an observation back to the model. The companion question on information fidelity covers the outbound path.
Checkpoint one of four
Is measurement reliable?
Noise, instrument drift, batch effects, failed runs.
AssessmentStandardized workflows can improve throughput and comparability in specific tasks, but do not establish that reproducibility is solved across instruments, materials and laboratories.
Confidence: high · 1 key entry
Key evidence
SAPP / DMX
→ Checkpoint clearedPeer reviewed · Public artifacts
2026.05
What it does
Standardises parallel protein production and measurement.
Bearing on this question
Batch production and SEC characterisation complete in roughly 48 hours with standardised output — the metrology rung is an engineering problem with a known solution shape.
Does not establish
It does not show feedback automatically improving the next design round; clearing this checkpoint says nothing about the three that follow.
Open evidence →
Further evidence
1 entry · corroborating
2026.08
Radical AI interview — PhD metallurgists annotate SEM images to capture what the company calls scientific intuition Testimony
Cleared
Checkpoint two of four
Does the metric match the research goal?
Whether the quantity being optimised is the quantity that matters.
AssessmentThe thinnest checkpoint. Almost nobody is studying it. A reward that covers a local metric can be optimised without solving the target, and the failure is invisible from inside the loop.
Confidence: low — the evidence is not negative, it is absent · 2 key entries
Key evidence
Lila Sciences — reward hacking in a physical verifier
→ Checkpoint failedExpert testimony
2026.07
Bearing on this question
An electrocatalyst can perform well in a short measurement and corrode quickly afterwards. If the reward covers only the local metric, the model will optimise it without solving the actual objective. This is the clearest public statement of the failure mode — and it comes from a team building the system, not from a critic.
Does not establish
Testimony, not a study. No public work quantifies how often this happens or how it would be detected.
Open interview →
Moderna V940
→ RefinesProspective · Clinical · Company-reported
2025.12
Bearing on this question
Phase 3 topline results support clinical benefit for the whole treatment system in one indication — the endpoint is unambiguously the right one. But the algorithm's isolated contribution is unknown, which separates two things usually conflated: measuring the right thing, and being able to attribute the result.
Does not establish
Effect size attributable to selection, and overall survival, remain undisclosed.
Open evidence →
Checkpoint three of four
Does the result change the model or experiment choice?
Whether the loop is closed in software, or closed by a person reading a report.
AssessmentArchitecturally described, never independently verified. Several systems publish the architecture; none publishes reward quality, calibration or cross-lab reliability.
Confidence: moderate on existence, low on quality · 2 key entries
Key evidence
Lila Sciences
→ Checkpoint clearedOperating evidence · Company-reported
2026.07
What it does
Treats the laboratory as a verifiable reward environment; lab instruments appear as tool calls inside the model's chain of thought.
Bearing on this question
The most explicit architecture for this checkpoint currently public, and the one that defines the loop as a training mechanism rather than a workflow.
Does not establish
Reward quality, calibration, cross-lab reliability and independent performance data are all missing.
Open interview →
Radical AI — campaign-based active learning
→ Checkpoint clearedOperating evidence · Company-reported
2026.08
Bearing on this question
Seven to ten campaigns run in parallel with results refreshed daily or every other day, and negative results are treated as training data rather than waste. Notably the loop is closed per campaign, not per experiment — a design choice that trades feedback latency for signal quality.
Does not establish
Synthesis still requires human judgement; no independent audit of the loop's operation exists.
Open interview →
Further evidence
1 entry · corroborating
2026.02
CuspAI interview — agents orchestrate computation and experiment, described as being at "various stages of maturity" Testimony
Cleared
Checkpoint four of four
Does the updated policy deliver repeatable gains?
Whether the feedback demonstrably made the following decision better.
AssessmentThis is the step most in need of validation. The cases on this page do not provide a complete controlled chain from observation to policy change and multi-round benefit. Evaluation should compare sample efficiency, decision quality and cumulative gain rather than require monotonic improvement.
Confidence: low · 1 key entry · the core finding of this page
Key evidence
HEA05
→ Closest availableProspective · Physical · Peer reviewed
2026.08
What it does
Four adaptive rounds select and fabricate new alloys, with eight new candidates prospectively tested.
Bearing on this question
The nearest thing anyone has published to a closed loop with a demonstrated round-over-round effect. Four rounds is enough to show the mechanism runs; it is not enough to show it compounds.
Does not establish
One alloy family. No cross-family, cross-lab or high-throughput replication, and no control arm selecting by other means.
Open evidence →
Counter-evidence and cross-pressure
What role should people retain inside the loop?
Recorded because a tracked question with one-sided evidence is an assertion. Two of these challenge the feasibility of closing the loop; the third challenges whether a fully closed loop is the right objective at all.
Max Welling — against the dark lab
→ Challenges the goalExpert testimony
2026.02
The claim
A completely automated laboratory — close the door, say "find something interesting" — is explicitly not the objective, and will not be for a long time. He expects the real breakthrough to happen with many humans in the loop.
Why it matters here
This is not a claim that closing the loop is hard; it is a claim that measuring progress by loop closure is the wrong frame. If he is right, the fourth checkpoint above should be rewritten to ask whether the loop makes the expert faster, not whether it removes them.
Open interview →
Lila Sciences — false positives cut both ways
→ Complicates checkpoint fourCompany self-report
2026.07
The claim
A false positive is a disappointment for the experimentalist but useful to the model, because it reduces uncertainty.
Why it matters here
It means the human intuition for "this round went badly" is not a reliable read on whether the loop learned. Any evaluation at checkpoint four has to be model-side, not morale-side.
Open interview →
AI4S.fyi synthesis · Current gap
We can produce measurements faster and more comparably than before, but rarely show that they reliably improve the next decision. Not one public system publishes the counterfactual.
What would change our view
- Checkpoint 2A pre-registered comparison of a proxy measurement against the true endpoint on the same samples, reporting how often they disagree.
- Checkpoint 3Cross-batch and cross-instrument error and calibration reports; failed-run and human-intervention logs.
- Checkpoint 4One traceable chain — observation, changed decision, improved next round — with a control arm selecting by other means.
- CounterA demonstration that a loop with humans kept in the decision seat outperforms one that removes them, which would reframe the question rather than answer it.
Related reading
Lila Sciences on experiments as a training signal, and on what goes wrong when a physical reward can be gamed.
The companion question. That page covers the outbound path from world to model; this one covers the return path.