Binding Computer-Use Approvals to Specific Actions
An approval that says "yes" to the wrong action is worse than no approval at all. Here is how Ethen's computer-use path ties each decision to one exact action.
An approval that says "yes" to the wrong action is worse than no approval at all. Here is how Ethen's computer-use path ties each decision to one exact action.
A browser agent proposes dozens of small operations in a single run: click this, type that, scroll, wait, read the page. When a human approves one of them, the security question is not whether the click happened but what exactly was approved. If the approval is a vague blessing — "yes, keep going" — then anything the agent does afterward inherits trust it never earned. A computer use action approval has to be narrower than that: bound to a specific action, inside a specific run, for a specific attempt, under a specific policy. Otherwise replay, substitution, and scope creep turn the approval checkpoint into theater.
This article walks through the approval-binding mechanism in Ethen's computer-use runtime: the code that converts a proposed action into stable bytes, requests an approval against those bytes, and then re-checks them before the action is authorized to run. It also states the limits plainly, because this mechanism is one control in a larger system, not a certification of anything beyond its own scope. Computer is a Platform-only product direction, and the repository audit this article draws on is explicit that the inspected code does not implement or certify canonical V3 ACT. Approvals are a mechanism. They are not proof of universal prompt-injection resistance, and this article makes no launch or availability claim.
Why approvals bind to bytes, not objects
The heart of the mechanism is a small function, canonicalComputerUseActionBytes, in the computer-use runtime's approval-binding module. It takes a ComputerAction and returns a Uint8Array: the canonical identity of that action, encoded as UTF-8 bytes. The platform approval service hashes those bytes when it records and later verifies the approval.
The comment above the function explains the design choice directly: do not hash a raw JavaScript object, because property insertion order is not a stable security boundary. Two objects that mean the same action could serialize differently depending on the order their properties were inserted, and two objects that serialize the same way could mean different actions if the serialization drops fields. Either failure breaks binding. By hashing bytes derived from a canonical identity string instead, the mechanism gets a stable input: the same executable action always produces the same bytes, and a different action produces different bytes.
The same comment adds a scope note worth quoting in spirit: the legacy canonical identity deliberately includes only executable action fields. That "deliberately" matters. Metadata about an action — display labels, planner commentary, surrounding context — can change without changing what the action will do. If those fields were part of the hashed identity, harmless rephrasing would invalidate approvals, and reviewers would learn to re-approve reflexively. Binding only executable fields keeps the approval decision attached to what will actually execute.
This is the first of four bindings the mechanism layers together: the approval is bound to the action's canonical bytes. The remaining three — run scope, attempt scope, and policy — answer the questions "where does this approval count?", "how many times?", and "under which rules?"
Requesting an approval: run scope and identity
The requestComputerUseApproval function assembles everything the platform approval service needs to record a pending decision. Its inputs read like a checklist of the binding dimensions:
- Tenant and project scope:
organizationIdandprojectIdplace the approval inside the same ownership boundary as the rest of the run. An approval granted in one project cannot authorize an action in another, because the scope travels with the request. - Run scope:
runIdties the approval to one execution of the agent loop. Runs are the unit of work a browser session performs; binding approvals to a run means a "yes" from Tuesday's session cannot be replayed into Thursday's. - Requester identity:
requesterIdrecords who or what asked for the decision. - Action bytes:
canonicalComputerUseActionBytes(input.action)freezes the proposed action into the stable byte form described above. - Policy binding:
policy, anApprovalPolicyBinding, names the policy under which the approval is requested. - Attempt-scoped scope: the caller's
ApprovalScopewith the current attempt ID folded into its constraints. - Expiry:
expiresAtbounds how long the pending approval stays valid.
Each of these fields narrows the approval. Remove any one of them and a class of misuse opens up: without tenant scope, cross-project replay; without run scope, cross-session replay; without action bytes, substitution of a different action under an old "yes"; without expiry, an approval that lives forever. The request path is where the mechanism declares, up front, exactly what is being asked.
Note what this function does not do. It does not evaluate whether the action is safe — that is policy's job, applied by the approval service and the surrounding permission checks. It does not execute anything. It packages a precise question — "may this exact action run, here, now, under this policy?" — and hands it to the service that records the answer.
Attempt binding: why one approval covers one try
Inside the request, a helper called withAttemptScope deserves attention out of proportion to its size. It copies the approval scope and adds one constraint: computerUseAttemptId, set to the current attempt ID. Both the request path and the authorization path apply it, so the attempt that was approved must be the attempt that runs.
Attempts exist because agent loops retry. An action can fail at the transport layer, time out, or return an ambiguous result, and the loop will often propose the same logical step again. Without attempt binding, a single approval could silently cover every retry of an action — including retries the reviewer never saw, under page conditions that changed since the approval was granted. With attempt binding, each attempt carries its own identity in the scope constraints, so a fresh attempt needs a fresh decision (or an explicit policy saying otherwise).
This is also a replay control. If an attacker — or a confused planner — captures an approval identifier and tries to reuse it for a later attempt, the scope check fails: the stored attempt ID does not match the presented one. The approval is single-use in a meaningful sense, not just marked consumed after the fact but bound to an attempt that cannot recur. The repository audit confirms this property in its test mapping: the canonical-approval-binding behavioral suite covers a bound action with one-shot consumption, as part of a small set of isolated unit suites (20 assertions passing) that also cover tenant scoping, worker lifecycle, SSRF address rules, and lexical research checks.
There is a subtlety here worth stating honestly. Attempt binding prevents an old approval from authorizing a new attempt, but it does not by itself prevent the planner from proposing a consequential action twice. The audit flags the absence of a persisted unknown-effect barrier in the active loop as a separate gap: if an action's real-world effect is ambiguous after a transport failure, nothing in the loop blocks the planner from trying a similar consequential step again. Attempt-scoped approvals make each try require its own decision; they do not make the second try safe. That distinction belongs to the loop and reconciliation design, not to the approval mechanism, and the article returns to it in the limitations section.
Authorizing the action: re-deriving everything
Requesting an approval and authorizing against it are two separate calls, and the separation is load-bearing. authorizeComputerUseAction takes the approval ID, the accessor's access scope, the run and attempt IDs, the action as currently proposed, the policy, and the scope — and then re-derives the bindings from scratch:
- It recomputes the action bytes from the action object presented at authorization time, rather than trusting bytes stored or passed by the caller.
- It re-applies
withAttemptScope, folding the presented attempt ID into the presented scope. - It passes the policy binding, the run ID, and the access scope through to the platform service's
authorizecheck.
Re-derivation is what makes the binding a check rather than a label. If the action object changed between request and authorization — a different target, different text to type, different coordinates — the recomputed bytes differ from the bytes recorded at request time, and authorization fails. If the attempt ID changed, the scope differs, and authorization fails. The authorize path trusts nothing about the action except what it can recompute and compare.
The access parameter adds one more dimension: who is asking to consume the approval must hold an appropriate access scope. And approvalId names the specific recorded decision being consumed, which is what lets the service enforce one-shot semantics — an approval, once consumed, cannot authorize a second action even if all the bytes match.
Read the two functions together and the protocol is clear: request freezes a precise proposal into the approval record; authorize re-freezes the live proposal and demands equality. Anything that drifted in between — action content, attempt, run, policy, accessor — breaks the match.
Policy binding: the rules travel with the decision
Both functions take an ApprovalPolicyBinding and pass it to the platform service. The policy binding names the rules under which the approval was requested and must be evaluated: which policy version or snapshot applies, and by implication what the reviewer (human or automated) was actually enforcing when they said yes.
Binding the policy matters because policies change. A team might tighten its rules after an incident — new domain restrictions, new sensitive-action definitions, shorter expiries. If approvals floated free of the policy that produced them, an approval granted under loose rules could authorize an action after the rules tightened. By carrying the binding through both request and authorize, the mechanism lets the service detect the mismatch: the recorded policy is part of what must match.
Here the audit adds a pointed limitation. The computer-use runtime also contains a platform-approval bridge that mirrors local human approvals into the platform service, and the bridge's policy snapshot is fixed rather than a live version of all policy facets. The audit's recommendation is to preserve the strong one-use claim machinery while replacing optional mirroring with one authoritative lifecycle. In other words: the binding shape is right — policy travels with the decision — but the surrounding lifecycle still has two sources of truth (local approval state plus an optional platform mirror) where it should eventually have one.
The expiry story carries a similar caveat. The local store uses a five-minute expiry while the bridge requests a one-hour envelope, and the audit notes that neither number is a measured V3 TTL. Expiry bounds are present, which is good; their exact values are engineering choices awaiting empirical grounding, which the audit states rather than hides.
What surrounds the mechanism: policy checks and their limits
Approval binding does not operate alone. Before an action ever reaches the approval checkpoint, the runtime's policy module combines action allowlists, domain checks, and risk heuristics to decide what needs approval and what is refused outright. Understanding that layer is necessary to avoid over-claiming what approvals prove.
The audit describes this policy layer precisely: it is not a data-flow authority. There is no conjunction of label, source, destination, purpose, and transform that could confine, say, a Gmail-to-CRM copy to exactly one approved field. A model-supplied target label can influence sensitive-action recognition, and page-instruction pattern matching counts as defense in depth only. These are honest boundaries. They mean the approval mechanism binds decisions strongly, but the classification feeding into those decisions — "is this action sensitive?", "does this page content constitute an instruction?" — remains heuristic.
This is why the claim limit on this article exists, and why it matters beyond editorial caution. A reader who learns that approvals bind to canonical bytes might conclude the system resists prompt injection: after all, injected page content cannot forge the bytes. That conclusion would be wrong in an important way. Injected content does not need to forge bytes; it needs to influence which action the planner proposes, or how the policy layer classifies it. The bytes then faithfully bind the approval to the wrong action. Strong binding makes each decision precise. It does not make each proposal trustworthy. The audit's broader findings — no verified data-flow enforcement, no mechanical egress gate on page text reaching the planner, planner self-completion paths — all sit outside the approval mechanism's scope, and none of them are fixed by it.
What the audit does and does not establish
Because this article leans on the September 2026 browser reconciliation audit for context, its scope deserves a short summary. The audit inspected runtime contracts, adapters, action execution, planner transport, loops, policy, approval binding, storage, worker lifecycle, UI, and route guards, tracing shared approvals, missions, routing, and research infrastructure. It ran five isolated unit suites — 20 assertions, all passing — covering tenant scoping, sequential worker lifecycle, one-shot action approval, SSRF address rules, and lexical research verification. It explicitly did not exercise a production database migration, a native browser, a hosted billable session, a real-world effect, or a live account.
Its headline outcome is the guardrail for everything written here: the inspected repository contains a reusable browser-agent foundation, but it does not implement or certify canonical V3 ACT. The approval-binding code is called out as a genuine strength — stable action bytes, attempt and run binding, policy and scope carried through platform authorization — inside a system with documented gaps: duplicated approval truth, fixed policy snapshots, missing browser-epoch and price re-derivation, takeover implemented as a status toggle, and verification heuristics that accept completion too readily. A dated report proves only its scope, and this article treats the audit as exactly that: evidence for the mechanism's shape and the system's limits as of the audit date, not a standing certificate.
Putting it together
The computer use action approval, as implemented in the inspected code, is a four-part binding:
- Action bytes freeze the executable content of the proposed action into stable, hashable form, excluding volatile metadata and avoiding the instability of raw object serialization.
- Run scope — organization, project, and run IDs — confines the approval to one execution context, blocking cross-project and cross-session replay.
- Attempt scope folds a per-try identity into the approval constraints and is re-checked at authorization, so each attempt needs its own decision.
- Policy binding carries the governing rules with the decision through both request and authorize, so approvals cannot silently outlive the policy that produced them.
Request and authorize are separate calls with re-derivation at the second step, which turns the binding from a label into a live check. The platform service hashes the bytes, records the decision, and enforces one-shot consumption — and the behavioral test suite pins the bound-action, one-shot property in place.
The honest ending is the one the evidence supports. This is a well-shaped mechanism for making approvals precise: it answers "which action, where, how many times, under which rules?" with code rather than convention. It does not answer whether the proposed action was wise, whether the page content that shaped it was trustworthy, or whether the surrounding system meets a broader action-certification target. Those questions belong to policy classification, data-flow control, loop design, and empirical testing — areas where the audit documents real foundations and real gaps side by side. Bind approvals tightly, and then keep working on everything around them. That is what the code shows, and what it does not yet show.