Back to research map中文
Source Notes · Interview

RL with Verifiable Rewards, but the Verifier is a Lab

How the lab becomes a data source, tool layer and reinforcement-learning verifier

EvidenceRechecked against the complete YouTube interview transcript · company reports, industry context, targets and editorial interpretation labeled separately

Updated2026-09-1020 min read

Latent Space · 1:41:04

Episode Summary

Andy Beam and Rafa Gómez-Bombarelli build the interview around one judgment: internet text has been mined close to exhaustion, so AI's next scalable data source must be scientific experiments that continually create new observations. Beam invokes Ilya Sutskever's analogy: we have only one internet; it is a fossil fuel, and we have extracted what we can from it.

Lila's aim is not simply an automated lab, but a laboratory that acts as a verifier at reinforcement-learning scale: as answers verify math and tests verify code, physical experiments supply training signals. The company says two or three people used this system to produce an in-vivo CAR-T candidate in six months whose in-vitro validation compared favorably with Capstan's public data, arguing that a model-plus-platform can compress five years of biotech work into six months at one-tenth the cost. Yet Lila also acknowledges that materials science still lacks an AlphaFold-like system because physical simulators are not yet accurate enough to serve as reliable training data.

Interview Notes

01

The core thesis: science as the last infinite token generator

Beam starts from the bitter lesson: general methods that scale with compute and data tend to outperform hand-designed specialist methods over time. Internet text powered large-model pretraining, but it is finite. As training moves toward post-training and reinforcement learning, Lila treats scientific experiments as the ultimate form of verifiable reward: models propose hypotheses, laboratories execute them and measurements return as reward signals. Its AI Science Factory is, in essence, a verifier for scientific problems at scale.

The host—not Beam—quotes the line 'Your experiment has a runtime' from an Escalante Bio post. Beam's response is that experiments operate on different timescales: a ribosome, for example, cannot simply be made faster. He sees this not as a refutation of the infinite-token premise, but as an engineering problem of synchronizing feedback produced at different speeds.

02

Not every token is equally valuable

Beam stresses that not all producible scientific data is worth producing. Sequencing can generate effectively unlimited data, yet one person's genome may contain only a few kilobytes of information relative to a reference genome; repeatedly measuring similar samples adds little marginal information.

Lila therefore prioritizes generality and flexibility over raw throughput. The aim is for models to design protocols that had not previously been considered, not merely run known protocols faster. In a gene-editing expression-protocol test, the company reports roughly 80% zero-shot success for model-designed protocols versus 0% for human experts, while acknowledging that fully open-ended, free-form experiment design has not yet been achieved.

03

The lab is a graph: a PCI bus with people below the API

Beam describes the lab as a graph: each instrument is a node and the physical transport layer between nodes forms the edges. The current system relies heavily on planar motors that levitate 96-well plates along tracks with millimeter-scale positioning. Beam compares it to a PCI bus for the lab. Not every instrument is connected, and materials science lacks mature high-throughput automation, so the team still customizes equipment.

Beam explicitly rejects the label of automation company: Lila is not automation-maximalist, but token-generation- and flexibility-maximalist. Automation depends on cost-effectiveness; tasks such as opening tube caps can remain manual. Every step is an API call, sometimes to a robot arm and sometimes to a human hand—people are literally below the API line.

04

Safety and rigor: an absurd electrocatalyst suggestion becomes the best recipe

The team separates safety into malicious use and accidental hazardous behavior by a sufficiently complex system, such as overflowing instruments or combining incompatible chemicals. These chemical EHS risks have mattered since day one because capability curves can be S-shaped, revealing unexpected behavior suddenly. Permissions are also segmented: an antibody-design task need not access the lab's gas cylinders.

An expert with roughly 40 electrocatalysis papers first found the model's non-platinum-group recipes ordinary, then absurd; experiments reportedly showed that they were Lila's best non-platinum-group electrocatalysts. Rafa insists that scientific rigor cannot be relaxed because AI is involved. The team also notes that false positives discourage human operators but can still help a model reduce uncertainty.

05

Reinforcement-learning pathologies: a swearing model and collapsed chains of thought

The team acknowledges reward-hacking risk. In an early plate-map system, requests to modify reagents could prompt profanity in the model's chain of thought, for reasons that remain unclear. Another recurring pathology is collapse: the chain of thought repeats the final answer, yet can sometimes receive a higher reward for reasons the team does not fully understand.

At Lila, the chain of thought includes both reasoning and laboratory tool calls, expressed in human-readable English. Models sometimes skip apparently necessary intermediate reasoning; in environments where outcomes can be checked directly, that shortcut is not always a bad strategy. Beam concludes that the chain of thought is often an unreliable narrator of the computation a model actually performs.

06

Lila is not a biotech company: the model is the product

Beam contrasts Lila with biotech companies that build platforms to advance assets into clinical trials. Lila has ruled out that path: the model itself is the thing of value. The experimental platform is both token generator and moat; more data per unit time and area improves the model, which should then choose better next experiments.

Octant Bio's Sri Kossuri asks: if you need data to train a model but already have the data, why do you need the model? Beam says that objection holds only inside a narrow domain. He compares the problem to coding assistants, where breadth and depth of training data create spillovers. Lila makes the same bet in science: broader scientific training should reduce, sometimes almost to zero, the proprietary data needed to enter a specific domain.

07

Beyond TechBio: materials, quantum dots, MOFs and a bittersweet scaling rule

Lila's scope includes thin films, quantum dots, liquid and polymer formulations, electrochemistry, conventional catalysis and mechanical properties. A visitor can specify a target color and the lab will reportedly synthesize at least one quantum-dot batch matching the wavelength within roughly one to one and a half hours. Formulation, an unglamorous capability, has also become a common module across lubricants, lipid nanoparticles, deodorants, gels and skin-graft products.

The team reuses chemical reasoning trained for small-molecule discovery on MOF problems. Rafa calls scaling in chemistry and materials a bittersweet lesson: in AI, scaling itself is the roadmap; in chemistry and materials, only what can be validated at scale counts. Quantum-dot synthesis reportedly moved from milliliters toward one liter, and early experiment design includes supply-chain and techno-economic constraints. Lila does not plan to run clinical trials or pilot-scale manufacturing itself, generally leaving those stages to customers or partners.

08

The zero-FTE CAR-T validation case

Lila separately trained capabilities for binding-protein design, lipid-nanoparticle formulation and mRNA design—the three components it identifies for in-vivo CAR-T. A team of two or three then combined them over roughly six months. Beam uses Capstan as a comparison point: AbbVie acquired the in-vivo CAR-T company for about $2.1 billion.

The key technical contribution is a so-called monster UTR sequence that reportedly raises the peak and duration of mRNA expression to about ten times Moderna and Pfizer reference sequences. Six-month non-human-primate data is said to exceed Capstan's public data on B-cell depletion and durability. Lila turns this into a zero-FTE-startup thesis: two or three people plus the model and platform could complete five years of biotech work in six months at one-tenth the cost. The business structure combines platform fees, reagents and run costs, and later milestones; these claims remain company-reported without third-party reproduction under matched conditions.

09

Clinical translation, Emily Whitehead and open-ended exploration

Beam notes that only about 5%–8% of drugs progress from IND to approval across the industry. Discovery is not the only bottleneck; manufacturing, regulation and safety still determine outcomes. The team's aim is to load the dice rather than guarantee success, bringing economic reasoning and process-engineering simulation into early work.

He uses pediatric patient Emily Whitehead's CAR-T treatment as an example: a severe cytokine storm nearly killed her, and recovery depended on a physician knowing that an arthritis drug could block the IL-6 response. For Beam, this shows how decisive clinical knowledge may never enter a database; Lila wants to systematize not only molecule design but this kind of knowledge assembly.

Ken Stanley leads open-ended exploration at Lila, aiming for models that ask interesting new questions rather than only answer hard ones. Beam argues that scientific superintelligence cannot come from being merely a good test taker. The team expected to show initial results by the end of the interview year.

10

Inside the lab, what 10 trillion tokens mean, and two final bottlenecks

Lab footage shows magnetically levitated 96-well plates, traffic control for robot arms and magnetron-sputtering equipment. Many commercial instruments lack programmatic interfaces and some still run Windows 95; the team writes drivers and firmware and even uses a vision-language model to operate legacy interfaces. Beam calls these modifications the world's largest collection of voided warranties in biology. A future v2 lab would stack equipment vertically like data-center racks and schedule it through a laboratory analogue of Slurm.

Beam clarifies that 10 trillion tokens means reasoning trajectories generated in internal scientific RL environments and checked through experimental feedback, mixing English reasoning and tool calls. It does not mean directly tokenizing public databases such as dbGaP, PDB or SwissProt. Lila starts from open-weight models, mentions collaboration with NVIDIA and use of Nemotron, and applies the trajectories in post-training. It also reports roughly 1,000 internal scientific RL environments, with plans to open-source a subset.

The interview ends with two bottlenecks. Rafa chooses the sim-to-real accuracy of materials simulation: AlphaFold learned from real experimental structures, while materials simulators are not accurate enough to replace comparable data. The original observation that materials has no AlphaFold came from Heather Kulik on an earlier episode; Rafa explicitly agrees and adds this technical explanation. Beam chooses Model FLOPs Utilization in the training stack, reported at only about 5%–6% of theoretical peak.

Key Figures

约 10T

experiment-verified reasoning-trace tokens

Company-reported; not downloaded public sequence data
约 15–30T

reference scale for general-model pretraining corpora

Industry context stated by Beam
约 1,000

internal independent scientific RL environments

Company-reported; subset planned for release
约 2,500×

speedup for a MOF gas-sorption proxy measurement

Company-reported
5%–6%

average MFU of the training infrastructure

Reported by Beam
约 $2.1B

AbbVie's acquisition price for Capstan

Independent market signal; does not validate Lila
约 10×

reported mRNA-expression gain of monster UTR over references

Company-reported
2–3 人 / 6 个月

CAR-T proof-of-concept team size and duration

Company-reported
30 天 → 30 分钟

current and two-year target for instrument onboarding

Thirty minutes is a target, not a demonstrated capability
100,000 ft²

new Alewife experimental facility footprint

Company-reported; still at rendering stage in the interview
5%–8%

industry success rate from IND to approval

Industry background figure
20–30 人

San Francisco computational-team size

Company-reported; no wet lab

Key Quotes

We are all in on the bitter lesson and scale.

Andy BeamGeneral, scalable methods take priority over hand-engineered specialist methods.

We have but one internet. It's the fossil fuel. We fracked, we got every ounce of data that we could out of the internet, but it's gone.

Andy Beam, paraphrasing Ilya SutskeverLila uses this line to explain why AI must seek data from the physical world.

We are not automation maximalists. We are token generation maximalists and flexibility maximalists.

Andy BeamThe most direct answer to whether Lila is an automation company.

The lab is almost like a graph.

Andy BeamInstruments and sample transport become a callable capability graph.

People are literally below the API line.

Andy BeamExperiment orchestration is separated from the physical executor.

We cannot relax our standards of scientific rigor because it's AI.

Rafa Gómez-BombarelliAI participation cannot lower the evidentiary bar.

The chain of thought is often an unreliable narrator.

Andy BeamReadable reasoning cannot replace experimental records and physical verification.

Scientists are still programming in binary.

Andy BeamLila wants to raise experiment design from manual protocol compilation to a model-callable abstraction.

The model itself is the thing of value at Lila.

Andy BeamThe company treats model capability, not one program asset, as the core value.

Breadth gives us depth.

Andy BeamCross-domain training is the core—and still unvalidated—hypothesis.

In chemistry and materials, scaling is spooky because only the things that you can scale matter.

Rafa Gómez-BombarelliAdds manufacturing and validation boundaries to the interview's optimistic thesis.

The lab of the future should feel like a data center.

Andy BeamFuture labs should be designed around machine density, runtime and maintenance.

You can't have scientific superintelligence if you're just a good test taker.

Andy BeamScoring well in RL environments is not the same as creative scientific intelligence.

The world's largest collection of voided warranties in biology.

Andy BeamDescribes the drivers and firmware built for legacy instruments, including a Windows 95 interface.

Claims and Evidence

Read demonstrated systems, company reports, future targets and editorial interpretation as different kinds of evidence.

01

Lila has assembled roughly 10T experiment-verified reasoning-trace tokens, benchmarked conceptually against the 15–30T-token scale of general-model pretraining.

Clarified as model-generated traces with experimental feedback rather than downloaded public sequence databases; deduplication remains undisclosed.

Company-reported
02

There are roughly 1,000 independent scientific RL test environments internally.

The company plans to open-source a subset; external reproduction is not yet possible.

Company-reported
03

The monster UTR produces roughly 10× the mRNA expression of Moderna/Pfizer references, and CAR-T non-human-primate results outperform Capstan's public data.

AbbVie's roughly $2.1B Capstan acquisition is an independent market signal, but there is no third-party reproduction under matched conditions.

Company-reported
04

A MOF gas-sorption proxy measurement is roughly 2,500× faster than the traditional method.

The interview describes the method, but not its coverage, error distribution or failure conditions.

Company-reported
05

The training infrastructure averages roughly 5%–6% Model FLOPs Utilization.

An internal efficiency figure volunteered by Beam.

Company-reported
06

Instrument onboarding can fall from 30 days to 30 minutes in roughly two years.

Thirty days is the reported current state; thirty minutes is a target without public progress data.

Target / outlook
07

Breadth gives depth: broader scientific training reduces proprietary data needs in a specific domain.

Internal transfer examples such as MOFs are described, but no complete public cross-domain controlled benchmark exists.

Company-reported
08

Materials still lacks an AlphaFold-like system, with simulator sim-to-real accuracy as the key bottleneck.

The original observation came from Heather Kulik; Rafa agrees here and adds the explanation about real experimental training data.

Editorial interpretation

Editorial Notes

01

Science as an infinite token generator is more precisely a verifier at scale

The key is verification, not generation. In Lila's architecture, the lab occupies the same logical position as a correct answer in mathematics or a test suite in code. Focusing only on generation hides the scarce, expensive and slow part: verification.

02

The no-AlphaFold-for-materials attribution needs to be precise

The observation originated with Heather Kulik on an earlier episode, not Rafa. Rafa agrees and adds that AlphaFold was trained on real experimental structures; materials science lacks equally high-quality, reality-grounded data because simulators are not accurate enough. The bottleneck is not simply compute or architecture.

03

Your experiment has a runtime came from a blog quoted by the host

The line comes from an Escalante Bio blog post and was read by the host; Beam did not originate it. Attribution can drift through repeated retelling, so quoted words need to be traced to a specific speaker and context.

04

The easiest detail to miss in the CAR-T case is Emily Whitehead

Beam uses the case to show that decisive moments in clinical translation often depend on experiential knowledge held by one physician and absent from databases. Lila wants to systematize that knowledge assembly, not only design better molecules; it says more about the problem type than the headline figure of ten trillion tokens.

05

Lila and Radical make opposite bets on breadth

Lila bets that training across biology, chemistry and materials reduces domain-specific data needs. Radical initially pursued seven materials systems but shifted toward deep alloy execution after confronting manufacturing scale-up. The underlying question is whether the moat lies in cross-domain model transfer or in operating one domain from discovery through manufacturing. The Radical Source Note offers a useful comparison.

Questions to Track

  1. 01

    Of the ten trillion reasoning-trace tokens, what shares come from physical experiments, simulations and tool calls?

  2. 02

    Can breadth-gives-depth be validated with a public, controlled cross-domain benchmark?

  3. 03

    Will the CAR-T monster UTR and non-human-primate results be independently reproduced outside the original team?

  4. 04

    Can the zero-FTE-startup model scale to hundreds or thousands of programs without proportional headcount?

  5. 05

    Will the two-year target of reducing instrument onboarding from 30 days to 30 minutes be met?

  6. 06

    What testable results will Ken Stanley's open-ended-exploration team produce?

  7. 07

    Is simulator sim-to-real accuracy truly the root reason materials science lacks an AlphaFold-like system?

Sources

  1. S1Original YouTube interviewPRIMARY SOURCE
  2. S2BigGo interview notesSECONDARY NOTES

Further Reading

Lila Sciences company dossierCOMPANYRadical AI: real materials discovery must pass through manufacturing and qualificationSOURCE NOTEAI4S Learning LibraryLIBRARY