AI Labs Need a “Supervisor of Record” Before Agents Touch the Physical World
Opinion | Guest Contributor
By Gleb Tsipursky, PhD
AI Labs Need a Supervisor of Record for Physical AI Agents
AI agents are crossing an important boundary. Anthropic’s Model Hardware Standard research preview lets agents operate physical devices including microscopes, liquid handlers and robotic arms, coordinate multiple instruments, adjust parameters in real time and, in some cases, recover from hardware errors without intervention. The same architecture is aimed at scientific labs and advanced manufacturing. That is a meaningful advance. It also exposes a governance weakness that has been easier to ignore when agents stayed inside software.
“Human in the loop” is too vague once an AI system can move a robotic arm, change a laser setting, transfer a liquid sample or trigger a manufacturing step. A human can technically exist somewhere in the process while having no clear responsibility, no defined intervention threshold and no practical way to reconstruct what happened after a failure.
Organizations deploying physical AI need something more concrete: a supervisor of record for every autonomous workflow. The idea is familiar in other high-consequence settings. When a decision matters, accountability works best when one person has a clearly defined role, authority and duty to intervene. Physical AI needs the same structure before autonomy becomes routine.
The supervisor of record should not approve every low-risk action. That would erase much of the value of automation. Instead, the role should define and own five things before the run begins.
1. The action envelope
First, the action envelope. The team should specify which devices the agent may control, which parameters it may change, the acceptable ranges and the operations that remain off-limits. Anthropic’s own description of MHS includes device information and safety limits that can be encoded for the agent. Those technical constraints should map to an explicit human authorization.
2. Stop conditions
Second, stop conditions. Teams should decide in advance what forces the workflow to pause. That could include an unexpected measurement, a device state outside a known range, repeated failed recovery attempts, a conflict between sensors or a step that would create an irreversible physical consequence. The point is to avoid inventing the escalation rule after something goes wrong.
3. Handoff thresholds
Third, handoff thresholds. Not every anomaly deserves the same response. Some can be handled automatically within a bounded recovery procedure. Others should route to a technician, scientist, engineer or safety lead. The supervisor of record should define those thresholds and make sure the people receiving escalations know they own the next decision.
4. Independent verification
Fourth, independent verification. An agent that performs an experiment should not automatically be the sole judge of whether the experiment succeeded. Consequential results should have a separate verification path appropriate to the domain, whether that means an independent measurement, a deterministic check, a second instrument, a human review or another validated procedure.
5. Reconstruction
Fifth, reconstruction. After a run, the organization should be able to answer what the agent observed, what it decided, which commands it issued, what the hardware did, when constraints were triggered and when a human intervened. NIST’s AI Risk Management Framework emphasizes governing, mapping, measuring and managing AI risk. Physical agents make that documentation discipline especially important because the consequences can leave the digital environment.
Anthropic’s early examples show why this structure matters. In one lab automation proof of concept, Claude coordinated a liquid handler, robotic arm and plate reader. In another example, the system adjusted a laser, observed the result through a camera and iterated before packaging what it learned into a deterministic script. This is precisely the kind of capability that makes autonomous experimentation attractive: the system can learn from feedback and turn successful behavior into repeatable execution.
But the same strength creates a new question. When an agent develops or modifies the procedure through exploration, who is responsible for deciding that the learned procedure is safe enough to reuse?

The answer cannot be “the model.” Responsibility belongs to the organization deploying it, and the supervisor-of-record mechanism makes that responsibility operational rather than rhetorical.
This approach should also make adoption faster. Engineers and researchers are more likely to trust autonomy when they know where the boundaries are, what happens when the system becomes uncertain and who has authority to stop the process. Clear guardrails create room for experimentation because people do not have to negotiate accountability from scratch every time.
The Financial Times described MHS as a step toward agents that can control real scientific and engineering equipment. That trajectory will extend beyond laboratories. Manufacturing, robotics, semiconductors and other physical systems are natural environments for agents that can reason across instruments and coordinate work. The governance model should move at the same speed.
Organizations should begin with low-consequence workflows, assign a supervisor of record, document the action envelope and stop conditions, rehearse one failure scenario, and review the resulting logs before expanding autonomy. The question after the pilot should not merely be whether the agent completed the task faster. It should be whether the team could tell when to trust it, when to interrupt it and who owned the decision. Physical AI will become more useful as agents gain access to more capable equipment. Before that access becomes ordinary, accountability needs a name attached to it.
About Dr. Gleb Tsipursky
Gleb Tsipursky, PhD, is a behavioral scientist, CEO of Disaster Avoidance Experts, and author of The Psychology of AI Adoption at Work: From Resistance to Results (Georgetown University Press, 2026). Learn more about the book.

