Research Infrastructure & Lab Automation· Evidence · Frontier

Anthropic Releases MHS: A Common Hardware Interface for AI Agents

MHS proposes a shared Interact layer for the AI4S loop: agents can discover, control and coordinate physical instruments through one interface, then return states, measurements and errors upstream.

Updated 10 min read

Research preview + prospective physical-device demonstrations

Current boundaryReal devices demonstrate orchestration, parameter adjustment, automated retries and control optimization, but not complete autonomous discovery; the standard remains a research preview and the evidence is primarily reported by Anthropic and its first partners

This work advancesLets AI agents operate and coordinate physical instruments through a common interface

  1. MODELUnderstand / predict
  2. DECIDEChoose next
  3. INTERACTAct / measure
  4. UPDATEChange next round

01 | What happened

How MHS works

On August 27, 2026, Anthropic opened a research preview of MHS, the Model Hardware Standard. The shared specification lets AI agents operate physical devices; early projects span microscopes, liquid handlers, plate readers, robotic arms and laser-control systems for quantum computers.[S1]

The basic path is scientific agent → MHS interface → device driver → instrument or robot, with state, measurements and errors returned upstream. Drivers expose common primitives such as read and write and generate a machine-readable reference file describing capabilities, adjustable parameters, current state and safety limits.[S1]

One useful mental model is MHS as an extension of MCP into the physical world. MCP gives models a common way to discover and call software tools; MHS applies a similar interface idea to instruments and robots. The analogy stops there, because physical devices must also handle live state, motion limits, hardware faults and safety interlocks.[S1]

Agents can control devices through MCP, a command-line interface or code. For low-latency or long-running work, an agent can package an operating sequence into deterministic code so the device executes without language-model reasoning at every time step.[S1]

MHS is model-agnostic rather than Claude-specific; any compatible agent harness can in principle use it. Access is currently limited to an initial group of research and manufacturing partners, with open sourcing planned after further physical-safety evaluation and deployment guidance.[S1]

Official Model Hardware Standard release image from Anthropic
Official Model Hardware Standard release image from Anthropic

The official visual accompanying Anthropic's MHS research preview. Image from Anthropic. [S1]

02 | What Changed

Which part of the scientific loop does MHS advance?

Its primary contribution is Interact. Once an upstream system decides what to do, MHS converts that decision into multi-instrument operations, reads device state and returns measurements and execution errors. It supplies neither a new scientific model nor the hypothesis or experiment-selection policy.[S1]

Some projects already form bounded loops of execute, measure, adjust and retry. Sustained scientific learning still requires an external model to interpret results, a decision layer to choose the next experiment, an update layer to change a model or policy, and safety and exception handling for long-running operation.[S1]

The narrower conclusion is that MHS reduces integration cost between Interact and Update, making autonomous loops easier to build without itself demonstrating autonomous scientific discovery.
Genentech experiment workflow orchestrating a liquid handler, robotic arm and microplate reader through Claude and MHS
Genentech experiment workflow orchestrating a liquid handler, robotic arm and microplate reader through Claude and MHS

Anthropic / Genentech: Claude coordinates three devices and reads their state through MHS; the orange path marks a bounded flow-rate optimization loop. Image from Anthropic's official release. [S1]

Why a common interface still needs device drivers

The analogy to USB for laboratory hardware is useful but incomplete. MHS does not eliminate drivers or make a manual instrument programmable. A device still needs an API, SDK, CLI, COM interface, file-drop system or automatable GUI, and its low-level differences still require a driver.[S1]

MHS standardizes how agents discover and call devices, not the devices' internal implementations. It translates capabilities, live state, operations and safety constraints into a discoverable interface for upstream systems.

From one-off glue code to a reusable device interface

Many instruments are already software-controllable. The hard part is that vendor software, programming languages, data formats and state representations differ. A new device often requires fresh glue code, protocol handling, state synchronization and recovery logic—work Anthropic and its partners say commonly takes weeks or months.[S1]

MHS attempts to replace one integration for every agent–instrument pair with one MHS integration per instrument that compatible agents and workflows can reuse. The immediate advance is interoperability, not stronger scientific reasoning.

Comparison of a traditional academic lab, a centralized automated lab and an MHS-based lab
Comparison of a traditional academic lab, a centralized automated lab and an MHS-based lab

Anthropic / UW Baker and Pinglay labs, Fig. 1: MHS connects distributed instruments through a common scheduling and monitoring layer while preserving research flexibility. Image from Anthropic's official release. [S1]

Model exploration, deterministic execution

MHS also demonstrates a useful division of labor for physical systems: a model explores a device, observes feedback and improves a control policy, then writes inspectable, testable deterministic code. QuEra used this pattern for laser control, with the final script running without the agent remaining online.[S1]

The language model handles open-ended exploration and improvement; deterministic code handles speed, stability and auditable execution. That is more realistic than keeping an LLM permanently inside a millisecond-scale control loop.

This division of labor improves execution efficiency and auditability; it does not supply physical understanding. A model can turn a useful procedure into code, but if it misdiagnoses a failure, deterministic execution may simply repeat the mistake more reliably. The Genentech case below exposes exactly that boundary.[S1]

03 | Key figures

From weeks of integration to 99.3% laser-relock success

7 → 1

Janelia microscopy moved from launching seven programs in sequence to one unified dashboard[S1]

Weeks → 8 hours

CMU built drivers, orchestration and an autonomous rerun from raw devices[S1]

R² < 0.9 → > 0.98

CMU's agent rejected a saturated curve, changed the concentration range and reran on a fresh plate[S1]

58% → 99.3%

QuEra laser-relock success after agent-guided optimization, measured across 700 blind tests[S1]

363 / 16 h / ≈10×

QuEra's unattended experiments, duration and reduction in residual error[S1]

6 devices / <1 week

UW Baker and Pinglay labs connected six devices, including driver work; a screening round can contain about 1,000 proteins, while one physical test costs roughly $100 and a week of labor[S1]

All figures above are partner-reported in Anthropic’s release and have not yet been evaluated under a common benchmark or independently replicated.

04 | Why it matters

The bottleneck is not instrument programmability, but coordination

The bottleneck is often not whether one instrument can be programmed, but whether different instruments can share state, respect dependencies and form one research workflow. If researchers still move results manually between vendor tools, even a strong experimental plan cannot continuously become physical evidence.[S1]

Traditional automation excels at repeating a fixed protocol thousands of times. Research changes samples, device combinations and conditions, and the next step may depend on the latest result. MHS places a reusable interface between deterministic equipment and changing research goals rather than replacing mature device control with an agent.

The economics at the UW Baker and Pinglay labs make the bottleneck concrete. Generating a protein design costs about $0.01, while physically testing one candidate costs roughly $100 and a week of labor—and one computational screen can produce about 1,000 candidates. The team connected six devices, including driver development, in under a week to narrow this widening gap between computational proposals and physical feedback.[S1]

Within design–build–test–learn, MHS primarily covers Build and Test plus device coordination and data handoff. Design still comes from scientific models or researchers, while Learn still requires an external model to interpret results, update strategy and choose another round.[S1]

Tetsuwan provides a larger-scale example of device characterization: the system ran 9,143 dispense tests, and its precision model improved on the manufacturer's specification by 12% on held-out conditions. A common interface can therefore do more than issue commands; it can return operational data to calibration and control workflows.[S1]

Partners announcing support, pilots or integration plans at launch included AWS (Strands Robots), Automata (LINQ), Danaher, Doosan Robotics, MBF Bioscience (ScanImage), QIAGEN (QIAsymphony Connect), Tecan (Fluent), Universal Robots, Hugging Face (LeRobot) and Raspberry Pi. That signals supply-side willingness to experiment with the standard; it is distinct from evidence that the same driver already works safely across laboratories and projects, which requires real deployments.[S1]

WHAT CHANGED

AI4S layer
Physical interaction / Laboratory hardware interfaces / Experimental automation
Original bottleneck
Instruments can run separately but expose incompatible interfaces, states and data formats, forcing each workflow to rebuild glue code
What changed
MHS connects agents and real devices through common drivers, device descriptions and MCP/CLI/code control
Key evidence
Real devices demonstrate multi-instrument orchestration, autonomous reruns, error recovery and control-script improvement across hundreds of experiments
Still unsolved
Cross-lab driver reuse, physical fault understanding, safe long-duration operation, broad adoption and complete autonomous discovery

EVIDENCE IN CONTEXT

Editorial status
Frontier
Evidence setting
Closed-loop or robotic lab
How the evidence was produced
Prospective · Physical devices · Bounded loops
Provenance and access
Company and partner-reported
Evidence ceiling
It shows that a common interface can support orchestration, reruns and control optimization on real devices, but does not establish transferable, long-running autonomous discovery.
Who did what
Humans define objectives and safety boundaries; agents orchestrate devices and adjust parameters in bounded tasks; external systems own scientific updates

05 | Evidence status and boundaries

What do the current demonstrations actually establish?

These are not software-only simulations. MHS has operated liquid handlers, robotic arms, plate readers, microscopes, lasers and quantum equipment. CMU and QuEra include prospective real-device runs, making the evidence stronger than an ordinary product demo or a simulation-only agent workflow.[S1]

The evidence remains first-party and partner-reported. The results appear inside Anthropic's official release rather than a shared benchmark, independent evaluation or peer-reviewed study. Devices, tasks, safety boundaries and metrics differ, so no single result transfers automatically to another lab.[S1]

CMU demonstrates bounded closed-loop execution. The system rejected a poor curve, changed the concentration range and reran the experiment, moving beyond fixed automation. The team also induced six failures—a missing plate, a rotated plate, a busy reader, a disconnected camera, an unreachable device and an emergency stop—and every one was blocked before robotic motion. But the experiment used a colorimetric dye in place of a drug, validating operational, safety and decision structure rather than autonomous drug discovery.[S1]

UW demonstrates lower-cost, faster laboratory integration. The Baker and Pinglay labs connected six distributed devices in under a week, including driver development, and linked a dashboard, qPCR stop recommendations and multi-device workflows. This is a strong demonstration of integration speed and real demand, but the same drivers have not yet been replicated across independent laboratories and different protein workflows.[S1]

QuEra demonstrates equipment-control optimization. The agent improved a control strategy through hundreds of real experiments and validated the script through blind tests and extended runs. The task had a defined objective, observable state and success criterion; it did not generate a new quantum-physics hypothesis.[S1]

A physical interface is not physical understanding. In Genentech's foaming example, Claude initially treated a physical problem as a software error and retried. A researcher had to explain the cause before it moved to a clean well and reduced mixing. The interface says what can be done and what the device reports; it does not explain why the physics failed.[S1]

Execution autonomy is not full scientific autonomy. Humans still define goals, candidate spaces, success criteria and exception boundaries. MHS supports substantial automated execution and local parameter updates, while scientific-model updates and the direction of the next research round remain external.

Fully manual legacy equipment still needs modification. MHS requires a software-operable entry point. Instruments that depend on manual knobs, manual loading or non-digital state still need vendor drivers, extra sensors, robotics or hardware modification.

06 | What I Learned

A physical interface is not physical understanding

Interact is an underappreciated layer of AI4SGenerating an experimental plan does not make the experiment happen. Plans must become device instructions, state must be observable and measurements must return in usable form. MHS frames that layer as an infrastructure problem that can be standardized.

Programmability is not the only bottleneckMany devices are already software-controllable; the problem is that they cannot share state or combine flexibly. Lowering cross-device and agent–device integration cost may matter more than adding another isolated automated instrument.

Execution autonomy is not decision autonomyAn agent can operate equipment with little intervention while humans define the objective, candidate space, success criterion and risk boundary. When a system runs unattended, we should still ask who chose the experiment and how its result changed the next decision.

Model exploration plus deterministic execution may be more practicalQuEra's pattern combines open-ended model exploration with the speed, stability and inspectability demanded by physical control. The model need not remain in the loop forever if it can produce a better deployable control policy.

An automated rerun is not yet scientific UpdateParameter adjustment can close a control loop, but a scientific loop must interpret what the result means and let new evidence change a model, hypothesis or experimental choice. MHS opens the return path; the upstream system still has to learn.

Sources

Factual claims link to original announcements, project lists, trial records or journal papers where possible. Research plans are kept separate from completed results.

  1. S1Previewing the Model Hardware StandardAnthropic · 2026.08.27 · Official research preview and partner case studies