Arco’s Intervention Threshold is measured by frequency: how often a task requires a human, roughly one intervention per hundred executions at T1. Reversibility enters the model at two points — the risk profile that mandates human involvement at T3, and the reversibility window that gates removing an approval step — but between those points, frequency is the variable that sets the threshold. That is right for almost everything a digital agent does, because a bad software decision costs a retry. It is the wrong variable for anything that cannot be retried, at any tier. The Physical Intervention Threshold is an Intervention Threshold calibrated by the reversibility of an action rather than by frequency alone — set by the cost of a single wrong action, not how often it occurs. Anthropic’s Model Hardware Standard just supplied the first public evidence the distinction is not theoretical.
What frequency was always measuring
Memo #104: What MTTI Doesn’t Measure made the first crack in the frequency assumption: MTTI counts how rarely a system needs help, and says nothing about how well it recovers once it does — which is why Arco tracks Rollback Cost alongside it, the cost of reversing an incorrect action once it is identified. That memo was written about software. A bad software decision has a Rollback Cost, but the floor on that cost is low: revert a commit, replay a transaction, refund a charge. Physical actions do not share that floor. A dropped sample, a mispriced batch of reagents, a laser fired at the wrong intensity — the cost of reversal is not low, and in some cases there is no reversal at all.
The obvious objection is that this just restates Memo #78: The Last Approval’s Authorization Trap — the organisational default to keep a human in the loop at consequential decision points, which Arco already treats as a failure to design away. It is not the same claim. The Authorization Trap describes reluctance: an operator who cannot bring themselves to remove an approval gate that no longer serves a function. The Physical Intervention Threshold describes a gate that should not be removed, because the task itself, not the operator’s nerve, is what sets the requirement.
The two variables the current model conflates
Frequency and severity are not the same axis, and between T1 and T3 Arco’s threshold model has scored only one of them. Escalation Rate sets how often a task actually escalates; the Intervention Threshold sets the frequency target Escalation Rate is measured against. Neither term asks what a single miss costs. A T1 task can carry a 1:100 Escalation Rate and a near-zero Rollback Cost — a wrongly categorised support ticket, reissued in seconds. A different task can carry the identical 1:100 Escalation Rate and a severe Rollback Cost — a produced batch that cannot be recalled, a reagent that cannot be un-mixed. The current model scores both tasks identically, because both clear the same frequency target. They are not identical, and a metric that cannot tell them apart is not measuring the thing that actually determines whether autonomy is safe to extend to that class of task in the first place.
What Anthropic’s own numbers show
MHS’s research-preview results make the distinction concrete rather than hypothetical, and they do it twice inside one deployment. Carnegie Mellon integrated a liquid handler, plate reader, robotic arm, and cameras in eight hours against a typical multi-week vendor setup, and at QuEra a laser-recovery routine improved from a 58% to a 99.3% success rate — both frequency-of-success improvements, exactly what a T1-style threshold rewards. The 99.3% figure came from 700 timed trials of a compiled recovery script, run with no agent in the loop at all: unattended execution of a finished, deterministic routine whose every action had already been tested.
The QuEra pilot’s account of the same deployment describes the other half of the record: an agent that “often stopped to wait for human confirmation before performing an action it deemed even slightly risky” — a near-constant intervention rate on new physical actions whose consequence the agent could not yet judge, inside a system otherwise operating at T1-grade reliability everywhere else. Read against a frequency-only threshold, that half looks like a failure: the system is escalating constantly. Read against a Physical Intervention Threshold, the two halves are one policy applied to two different Rollback Costs — release what is proven and reversible, hold what is new and not. The agent is not less capable on the risky actions. The actions are less forgiving. QuEra did not design that threshold; the agent’s default caution exposed where it should sit.
The Operator’s Verdict
A threshold tuned only for frequency will eventually approve something it should have stopped, because it was never asking the question that matters for an irreversible action. The fix is not a stricter number. It is a second variable: what does it cost to be wrong here, once, regardless of how rarely it happens.
An operator extending the Stewardship Model into physical execution should score every task on both axes before assigning a threshold — how often, and how expensive if wrong — rather than importing the digital default wholesale and finding out where the two axes diverge only after the divergence has already cost something.
Technology changes what is possible. Reversibility determines what is safe to automate.
KEY TAKEAWAY
What is the Physical Intervention Threshold, and how is it different from Arco’s standard Intervention Threshold?
The standard Intervention Threshold calibrates how often a task escalates to a human, targeting roughly 1:100 at T1. The Physical Intervention Threshold calibrates the same decision by the action’s reversibility instead — a high Rollback Cost requires a tight threshold regardless of how rarely it occurs, because what matters is the cost of one wrong action, not its frequency. MHS is the first public evidence that the distinction is real: one deployment ran a compiled recovery script unattended for 700 trials and paused for confirmation on nearly every new physically risky action. Key metric: MHS integration at Carnegie Mellon completed in 8 hours against a typical multi-week setup — frequency-grade performance — while the same deployment held human confirmation on nearly every physically risky action, an intervention rate that would fail a 1:100 target but is correct once reversibility, not frequency, sets the threshold.
