Paper Note· Evidence · Frontier
David Baker Team Paper: AI Can Design Thousands of Proteins. How Does the Lab Keep Up?
David Baker's team did not build another protein generator. It focused on a more physical AI4S bottleneck: models can produce candidates far faster than laboratories can validate them.
●Peer-reviewed paper + public research materials
Current boundary:it remains an IPD screening workflow, without validation on complex proteins, cross-lab replication or automated next-round selection
This work advances:Moves protein candidates into physical production and standardized measurement faster
- MODELUnderstand / predict
- DECIDEChoose next
- INTERACTAct / measure
- UPDATEChange next round
01 | What happened
On August 20, 2026, David Baker's team published “Accelerating protein design by scaling experimental characterization” in Nature Communications. The work comes mainly from the University of Washington Department of Biochemistry and Institute for Protein Design, with live RSV neutralization experiments performed by collaborators at Karolinska Institutet. It introduces two connected experimental workflows, SAPP and DMX.[S1]
The problem this paper addressesRFdiffusion and ProteinMPNN can propose thousands of candidates in days; a conventional wet lab may take weeks to express, purify and measure dozens. The faster design becomes, the longer the queue for physical feedback grows.[S1]
A sequence that looks promising in silico may still express poorly, aggregate or form the wrong oligomeric state. Those questions return to the lab. Large pooled assays can screen many sequences, but often read only a small number of properties such as binding or stability and do not reveal whether each individual protein is easy to produce, purify and recover in the intended state.[S1]
Baker's team did not build another design model. It redesigned the Build and Test stages after Design. The paper says SAPP and DMX have already been used inside IPD to produce and characterize tens of thousands of de novo designed proteins. Their job is not to identify a perfect answer in one shot, but to let the physical world answer more quickly: which designs actually work?[S1]
02 | What Changed
SAPP turns protein experiments into a batch workflow
Where the old queue stalledAfter DNA assembly, a conventional workflow still transforms, plates, picks a single colony, grows it, prepares plasmid, verifies the sequence and then moves the confirmed construct into an expression strain. This takes days and colony picking is difficult to parallelize at scale.[S1]
Engineering tradeoffSAPP uses Golden Gate Assembly and places the lethal ccdB gene in the destination vector. A correct insert replaces ccdB, while empty background vector is strongly suppressed. The assembly can therefore enter an E. coli expression strain directly. Moving 192 reactions from linear DNA to inoculated expression cultures takes about two hours.[S1][S4]
Is the shortcut reliable?The shortcut does not eliminate DNA synthesis and cloning errors. Across 929 assembly reactions, 89.3% had a dominant clone above 90% purity. Mass spectrometry of 863 purified proteins found about 90% within ±1 Da of the expected mass. The intended construct usually dominates enough for screening, at the cost of occasional false negatives.[S1]
The point is not that sequencing no longer matters. Screening and final validation are designed as separate stages: a faster workflow with some tolerance for missed positives first covers hundreds or thousands of candidates, while final hits still receive strict clonal and sequence confirmation.
The more important step: SEC for every sample turns experiments into a data line
Traditional 96-well screens often stop at SDS-PAGE to see what expressed, then scale up a few candidates. Expression is far from proving that a protein is good: it may aggregate, form the wrong oligomer or fail to preserve the intended native state.[S1]
SAPP sends every sample through affinity purification and then directly into size-exclusion chromatography (SEC). SEC is not only purification: each chromatogram reports soluble yield, dispersity, estimated molecular weight, aggregation and oligomeric state. A batch of 192 samples runs overnight in about 15 hours, under five minutes per sample.[S1]
A Python pipeline analyzes chromatograms in batches and generates liquid-handler instructions for fraction pooling and concentration normalization. The output is not only a 96-well plate of purified proteins but also a standardized file containing the experimental results. The full SAPP workflow takes about 48 hours, with approximately six hours of bench work.[S1][S3]
From an AI4S perspective, SAPP matters not only because experiments are faster, but because every design returns comparable, machine-readable physical data in the same format.

Excerpt from paper Fig. 1. Reading focus: start with the upper workflow from assembly through SEC and robotic pooling; the lower panels validate clonal purity, mass, yield and reproducibility. Qian et al., Nature Communications (2026), CC BY 4.0. [S1]
Why does throughput matter? Geometry is hard to rank in advance
A fluorescent-protein study first showed that SAPP could return comparable data for 96 redesigned constructs within a week. Some variants shared only 50–80% sequence identity with the wild type yet remained monomeric and retained absorbance after one hour at 95°C; others showed clear shifts in emission spectra.[S1]
The RSV study makes the case more clearly. The team screened about 31,000 minibinders targeting site III of RSV F, identified cb13 at 35 nM affinity and found by cryo-EM that the observed complex closely matched the design, with a binder-backbone Cα RMSD of 1.9 Å. They then placed cb13 on 27 oligomerization scaffolds to make 58 constructs. SAPP selected 19 with suitable expression and assembly for live-virus neutralization.[S1][S5]
Two constructs reached IC50 values of 40 pM and 59 pM, compared with 5.4 nM for monomeric cb13 and 156 pM for the MPE8 antibody control. The failures and branches are more informative: two cb13 dimers differed by about 25-fold; moving the binding domain from the N to C terminus on the same trimerization scaffold changed potency by nearly tenfold; and C6 or C8 architectures did not automatically improve activity.[S1]
Throughput therefore does more than validate a model's top few candidates faster. It lets researchers search a geometry space that is still difficult to rank reliably. When only three or five constructs can be made, this systematic exploration rarely happens. At dozens or hundreds, it becomes an experimentally answerable question.

Excerpt from paper Fig. 3. Reading focus: panel f shows roughly 25-fold differences in RSV neutralization when the same cb13 binder is placed in different oligomer geometries; the other panels cover design, assembly and structural validation. Qian et al., Nature Communications (2026), CC BY 4.0. [S1]
Once SAPP worked, the next bottleneck appeared: DNA cost
This is the paper's clearest systems story: bottlenecks move. Once protein expression, purification and initial characterization accelerated, DNA synthesis became the dominant cost of large campaigns. Ordering individual gene fragments is convenient, but at hundreds or thousands of designs DNA absorbs most of the budget.[S1]
Cheap DNA is pooled, but experiments need it arrayed: one construct per well, with a known identity.
Oligo pools are much cheaper, but all sequences arrive mixed together. DMX does not solve a new protein problem; it solves this format mismatch by turning a mixed pool back into a clone library that can enter experiments one known construct per well.[S1][S2]
DMX: one sequence per well, with its identity known
DMX clones a DNA library into a common vector and picks colonies into 384-well plates. Genes in each bacterial lysate receive combinatorial barcodes, are pooled again and sequenced with Nanopore long reads to map barcodes to gene sequences. An automated pipeline outputs a pick list, and a liquid handler re-arrays confirmed colonies into new plates.[S1][S3]
Four sets of 24 barcodes can theoretically identify 24⁴, or 331,776 wells. In a test library of 1,500 designs, the team analyzed 4,608 colonies and recovered 78% of target variants, approaching the roughly 92% theoretical limit at that sampling depth.[S1]
The paper estimates that DNA for a 500-design campaign is about fivefold cheaper than individual gene fragments, and an approximately 2,000-design campaign about eightfold cheaper. The cost is roughly five extra days. SAPP alone is simpler for one or two hundred designs; at thousand-design scale, DMX followed by SAPP becomes more attractive.[S1][S2]

Excerpt from paper Fig. 4. Reading focus: follow oligo pool → colony picking → combinatorial barcode → Nanopore mapping to see how DMX returns a mixed pool to an arrayed, sequence-identified clone library. Qian et al., Nature Communications (2026), CC BY 4.0. [S1]
03 | Key figures
One SAPP production and standardized-characterization cycle[S1]
Researcher bench time during the full cycle[S1]
SEC time per sample[S1]
Samples per chromatography instrument, already an emerging hardware boundary[S1]
Share of 929 assembly reactions with dominant-clone purity above 90%[S1]
Estimated DNA-cost advantage of DMX over individual gene fragments[S1]
04 | Why it matters
SAPP matters for more than speed. Each design passes through a comparable production and characterization workflow, with soluble yield, dispersity, estimated molecular weight, aggregation and oligomeric state organized into a common data structure.[S1][S3]
The wet lab begins to look like a data line that repeatedly returns machine-readable ground truth. Models propose candidates, the lab builds and measures them, and the resulting data can then feed a future round of model training or active learning.[S1]
More precisely, SAPP and DMX are experimental infrastructure required for a self-driving protein lab, not a completed self-driving lab. They advance Build, Test and Data, but do not yet automate selection of the next experiments from the previous round's results.[S1]
WHAT CHANGED
- AI4S layer
- Experimental validation / Lab infrastructure
- Original bottleneck
- Protein production + biochemical characterization
- What changed
- SAPP: complete a cycle in 48 hours, process hundreds of designs per day and return standardized experimental data
- Key evidence
- 48 hours | One SAPP production and standardized-characterization cycle
- Still unsolved
- Complex proteins, application-specific assays and full closed-loop experiment selection
EVIDENCE IN CONTEXT
- Editorial status
- Frontier
- Evidence setting
- Wet-lab validation
- How the evidence was produced
- Physical
- Provenance and access
- Peer reviewed · Public artifacts
- Evidence ceiling
- It shows that protein production and standardized characterization can be accelerated, but does not establish that feedback improves the next design round.
- Who did what
- Automation handles repetitive steps; researchers define candidates, supervise the workflow and run application-specific assays
05 | Evidence status and boundaries
A peer-reviewed methods paper, not proof of a complete closed loopThe paper provides workflows, scale measurements, case studies and public code, while experimental-data-informed design models and active learning remain future directions. It supports progress in experimental infrastructure, not a claim that autonomous protein discovery is already closed-loop.[S1][S3]
Screening, not final gold-standard validationSkipping single-clone sequencing creates false negatives, and polyclonality from DNA synthesis errors becomes more pronounced as genes get longer. Final hits still require strict clonal and sequence confirmation.[S1]
Applicability remains boundedThe workflow is optimized for engineered or designed proteins that express at useful levels. Proteins requiring complex expression conditions, special purification, other hosts or application-specific functional assays do not fit directly.[S1]
Evidence to watch nextThe next evidence should show whether labs without IPD's accumulated experience can reproduce the workflow, whether it extends to harder proteins and functional assays, and whether feeding standardized data back into models improves the next round's hit rate. Code, Addgene plasmids and the cb13 structure are public, and the authors disclose related patent applications.[S1][S3][S4][S5]
06 | What I Learned
01 | AI4S bottlenecks move; they do not disappearThe most memorable result is not 48 hours or 40 pM, but a systems rule. Faster design exposes wet-lab validation; faster validation exposes DNA cost; cheaper DNA then exposes SEC, functional assays and harder expression systems. When evaluating an AI4S model or company, a useful question is: which bottleneck in Design → Build → Test → Learn did it move?
02 | Screening and final validation need not use the same precision standardSAPP accepts some false negatives to cover more design space, then strictly confirms final hits. This does not remove rigor; it optimizes screening and confirmation for different jobs. Requiring final-validation certainty for every first-round candidate can exhaust time and budget before strong candidates are found.
03 | Good lab automation redesigns the workflow before adding robotsSAPP's largest gains come from removing steps that do not scale, standardizing plate formats, reducing intervention with auto-induction and combining purification with characterization. Robots take over repetitive work such as pooling and normalization. Lab automation does not mean robots everywhere; process design is often the first lever.
04 | Throughput changes which questions can be askedIn the RSV study, the same binder at the same valency differed by 25-fold when geometry changed. With only three or five constructs, systematic exploration is difficult. At dozens or hundreds, questions once left to intuition become experimentally searchable. Throughput does not only validate hypotheses faster; it expands the scientific design space that can be explored.
05 | Closed-loop AI needs standardized physical data, not only faster experimentsIf researchers must still inspect every image manually and each project uses a different format, faster experiments do not create a learning flywheel. SAPP's more important contribution is organizing each design's result into a common data structure. Design, Build, Test and Learn can connect only when the wet lab repeatedly returns this kind of ground-truth data at low cost.
Sources
Factual claims link to original announcements, project lists, trial records or journal papers where possible. Research plans are kept separate from completed results.
- S1Accelerating protein design by scaling experimental characterizationNature Communications · 2026.08.20 · Peer-reviewed paper
- S2Preprint with early data and code entry pointbioRxiv · 2025.08.06 · Preprint
- S3SAPP and DMX analysis and automation codeGitHub · bwicky/SAPP_DMX · 2025 · Open-source code
- S4Plasmids and barcode reagents for SAPP and DMXAddgene · 2025 · Research materials
- S5Structure of RSV F bound to the designed binder cb13RCSB Protein Data Bank · 2026 · Public structural data
