Skip to content

EthenEthenEthen

Research Note · 2026-10-03 · Trust & Accountable AI Work

Mandates: Compiling Human Intent Into Bounded Agent Authority

Publication type
Research Note
Evidence status
Proposal / Hypothesis: A proposed direction, system, or hypothesis that remains untested.architecture proposal; cites one published measured Ethen system card (AgentTrustBench R08)
Research program
Trust & Accountable AI Work
Published
Authors
Ethen Research Lab
Reading time
12 min read

An agent's authority today is scattered across prompts, API keys, role settings, approval rules and budget caps. We propose one object, the mandate, from which all of them are compiled.

Cover image for "Mandates: Compiling Human Intent Into Bounded Agent Authority". Decorative abstract motif; contains no data.

Abstract

When a person asks an AI agent to "clear this week's billing backlog, but check with me before refunding more than $200", they express a purpose, a scope, a budget, a risk threshold and an approval rule in one sentence. Current systems implement that sentence in pieces: a system prompt, an API credential, a role in an identity provider, an approval workflow and a spending cap, each configured separately and none aware of the others. This note proposes the mandate: a single, explicit object that records what a human principal has authorized an agent to do. It is then compiled into the enforcement artifacts that operate where the agent acts: attenuated capability tokens, constraints on allowed configurations, budget meters and approval routing. We describe the mandate's fields and its autonomy levels per action class. We describe delegation with attenuation, expiry and revocation, and the conditions under which AI agent authorization can be tested rather than assumed. We cite one published, narrowly scoped Ethen measurement of boundary enforcement. Everything else here is an architecture proposal.

The problem: authority with no single source

An agent that acts on behalf of a person or organization needs authority. In most deployments that authority is assembled from five mechanisms built independently:

  1. Instructions in the system prompt ("never refund more than $200").
  2. Credentials that grant access to tools and data, often broader than the task needs.
  3. Roles in an identity system, designed for humans.
  4. Approval workflows that pause for human confirmation, configured per tool.
  5. Budgets for model spend, often enforced at the account level rather than the task.

Each mechanism answers a different question, and none answers the one that matters: what exactly did this person authorize this agent to do, for how long, and under what conditions? The gaps between mechanisms are where failures occur. An instruction can be overridden by injected content. A credential can permit what the instruction forbids. An approval can be granted for one action and then applied to a modified version. A budget can be exhausted by a different task in the same account.

The first of these is the most serious. Prompt injection, in which untrusted content read by an agent alters its behavior, remains an open problem, and benchmarks such as AgentDojo exist precisely because defenses are incomplete (Debenedetti et al.). Policy-compliance benchmarks for web agents show that task success and policy adherence can diverge (Levy et al.). Emulated sandboxes have been used to surface risky agent actions before they reach real tools (Ruan et al.), which is useful for testing mandates but does not replace enforcing them. An authority model that lives in the prompt is an authority model that injected text can reach.

Proposal: the mandate

A mandate is a structured grant of authority from a principal to an agent. In the illustrative schema below, the principal is a human or an organizational role.

Mandate {
  mandate_id, principal, purpose            # who grants, and why
  scope { resources, systems, data_purposes }
  budget { money, tokens, wall_time, human_review_minutes }
  autonomy[action_class] ∈ { auto, notify, approve, forbidden }
  verifier_requirements                      # what evidence closes the task
  expiry, revocation_epoch
  delegation { max_depth ≤ 3, allowed_child_scopes }
}

The mandate is not enforced by the model. It is compiled (Figure 1) into four artifacts, each enforced by deterministic machinery at the point of action:

Left: a human principal grants a mandate containing purpose, scope, budget, autonomy levels per action class, verifier requirements, expiry, revocation epoch and delegation depth. A compiler box in the middle produces four outputs on the right: attenuated capability tokens at tool boundaries, routing and configuration constraints, budget meters, approval routing rules.

Figure 1. A mandate compiles into four enforcement artifacts. The human-facing mandate is one object. It compiles into enforcement artifacts that operate where the agent acts: attenuated capability tokens at tool boundaries, constraints on which execution configurations are allowed, budget meters, and approval routing. None of these is enforced by the model's prompt. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (Mandate); not implemented as described.

  • Capability tokens. Short-lived, audience-bound credentials carrying only the permissions the mandate allows for this task. Attenuation, the property that a holder can narrow but never widen a credential, has a long history in authorization design, including macaroons with contextual caveats (Birgisson et al.).
  • Configuration constraints. Limits on which models, providers, tools and sandboxes the agent's decision layer may choose. A mandate that forbids sending data outside a region forbids routing to a provider in another region. Action authority of this kind is distinct from data-use rights, which govern what may later be learned from the work; see Rights as Infrastructure. See Faros.
  • Budget meters. Reservations against money, tokens, time and human-review minutes, checked before each consequential step rather than reconciled afterward.
  • Approval routing. Rules that determine which action classes pause for whom, and with what second factor.

This framing has a practical consequence: the user-facing interaction ("give the agent a mandate"), the security model and the cost model become the same object. A person can read what they authorized. An enforcement layer can check it. A Work Receipt can cite it.

Relation to existing standards

The mandate is not a new protocol. It sits above existing authorization standards and compiles into them. OAuth 2.0 defines delegated access with scopes (RFC 6749). OAuth Token Exchange allows a token to be exchanged for another with different, typically narrower, properties, which suits multi-hop delegation (RFC 8693). OAuth Rich Authorization Requests express fine-grained, structured authorization details beyond flat scope strings (RFC 9396), and are perhaps the closest existing standard to the mandate's scope field. Workload identity frameworks such as SPIFFE provide service identities for the software that executes agents. A mandate system should map onto these standards rather than replace them. What the standards do not provide is the agent-level object: purpose, autonomy per action class, verifier requirements and a delegation policy, bound together and compiled consistently.

Autonomy levels per action class

Prompting a human for every step defeats the purpose of delegation and invites approval fatigue: people asked to confirm too often stop reading what they confirm. Human-factors research on automation describes related complacency and bias effects (Parasuraman & Manzey). Granting full autonomy is unacceptable for consequential actions. The mandate resolves this tension by setting autonomy per action class rather than per agent (Figure 2).

Matrix of seven action classes (read in scope, draft or internal reversible write, external message, reversible record update in scope, purchase or money movement, permission or policy change, destructive irreversible action) against four autonomy levels (automatic, notify, approve, forbidden to delegate). Money movement requires approval plus second factor; permission or policy change is never delegable to an agent.

Figure 2. Autonomy levels by action class: a proposed default. Autonomy is set per action class, not per agent. Irreversible or authority-changing actions are never automatic; reversible, in-scope writes can run under standing bounded authority. These defaults are a proposal from Ethen's internal architecture work, to be tested in usability and safety studies. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal; proposed defaults, untested.

The proposed defaults reflect one principle: strict where loss is irreversible, permissive where work is reversible and in scope. Money movement always requires explicit approval, with a stronger second factor for high-risk amounts. Permission, policy and key changes are never delegated to an agent at all. Reversible writes within an explicitly bounded scope, such as drafts, internal notes and bounded record updates, can proceed under standing authority with audit. These are defaults to be studied, not established findings. The relationship between autonomy settings, approval experience and outcomes is an open research question.

Approvals must bind to exact actions

An approval is meaningful only if it applies to the action that is actually executed. Mandate-based approval therefore binds to the action's exact parameters, the principal, the environment, the policy version in force, the resource scope, an expiry, and a hash of the material content. If the agent or the environment changes anything material after approval, the approval no longer applies and must be re-requested. This closes a common gap in which a harmless plan is approved and a different effect is executed after a model edit, a retry, or a change in the target's state. The interaction between approval binding and ambiguous outcomes is discussed in Unknown Effects in Autonomous AI Systems.

Delegation narrows authority

Agents increasingly delegate to subagents. Each delegation hop is an opportunity for authority to widen accidentally, through a confused deputy that uses its own broader credentials on a caller's behalf. The mandate's delegation rules (Figure 3) are designed to make widening structurally impossible.

Chain of four boxes from left to right: user mandate, agent (depth 1), subagent (depth 2), subagent (depth 3), each with a narrower scope. A fifth attempted hop at depth 4 is shown denied. Below, a box lists never-delegable permissions: identity administration, policy write, key export, impersonation, cross-tenant access, approval grant, residency override, evidence mutation, delegation widening.

Figure 3. Delegation narrows authority at every hop. Each subagent receives a distinct principal and a grant that is the intersection of its parent's grant and its own declared need. Depth is capped. A fixed list of permissions can never be delegated, whatever the mandate says. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (delegation attenuation; never-delegable list).

Every subagent receives a distinct principal and a grant equal to the intersection of its parent's grant and its own declared need. Delegation depth is capped; the internal proposal is three. A fixed list of permissions is never delegable whatever the mandate says: identity administration, policy writes, key export, impersonation, cross-tenant access, granting approvals, overriding data residency, mutating evidence, and widening delegation. No agent may approve its own privilege expansion.

Expiry and revocation

Authority that cannot be withdrawn is not bounded. Mandates expire. They can also be revoked at any time by the principal or an administrator. Token lifetimes alone cannot deliver fast revocation: a fifteen-minute token remains valid for up to fifteen minutes after the mandate behind it is withdrawn. The design therefore pairs short-lived tokens with an online authorization check immediately before each consequential action, so that revocation takes effect at the next action rather than at the next token refresh.

Ethen's internal architecture work proposes engineering targets for study [PROPOSED TARGET]: stopping an active agent within about five seconds of a kill request; denying new consequential actions within thirty seconds at the 99th percentile after revocation; token lifetimes of fifteen minutes or less; and authorization-epoch rotation within five minutes. These are targets for measurement under a defined threat model, not implemented service levels. Actions already in flight when revocation arrives raise the same ambiguity as any interrupted effect.

What has been measured

Ethen has published one narrowly scoped measurement relevant to mandate enforcement: the AgentTrustBench system card for run R08. On one pinned build, 197 enumerated conditions covering allow and deny decisions, stop conditions, spend limits, approval binding and recovery were each run once and then replayed in a different order. [MEASURED ETHEN RESULT] 194 of 197 conditions matched both the pre-registered prediction and an independent oracle on both passes. The three non-matching conditions are described in the card. As the card itself states, the result applies only to that build and that enumerated set. It supports no population claim, no product ranking and no claim about live traffic, and no statistical test was applied. It is cited here as an example of how mandate-style boundaries can be tested, not as evidence that the full design in this note works. The benchmark design that would test mandates broadly is VerifiedWork Control.

Mandates and commitments

A mandate states what an agent may do. It does not state what the agent must accomplish, or what remains unfinished. That is the role of a commitment graph: the record of obligations, preconditions and evidence that determines when work is actually complete. The two are complementary. Mandate violations are failures of authority; commitment violations are failures of completion. A system needs both to say, at the end of a task, "this was allowed, this was done, and this is how we know."

Failure modes and open questions

  • Compiling natural language is itself error-prone. A model that translates "clear the billing backlog" into a mandate can misread scope. The compiled mandate must be shown to the principal in a reviewable form before it takes effect.
  • Over-broad mandates. Principals may grant wide mandates for convenience, recreating the problem the mandate was meant to address. Interfaces that make scope legible, and defaults that start narrow, are design obligations. Neither is solved.
  • Action-class taxonomy drift. New tools introduce actions that do not fit existing classes. An unclassified action should default to the most restrictive class.
  • Approval fatigue. Even per-class autonomy may produce too many approvals in some workflows. Whether mandates reduce approvals per task at constant safety is an empirical question.
  • Usability. Users may find mandates harder to reason about than per-action prompts. That would argue for changing the design, not for weakening enforcement.

Limitations

This note describes an architecture that has not been implemented as described. The only Ethen measurement cited covers one pinned build and one enumerated condition set; it does not validate the design here. The proposed autonomy defaults and revocation targets are untested. The mandate addresses authority; it does not, by itself, defend against every way an agent can misuse authority it legitimately holds.

Conclusion

Agent authority should be something a person can read, a machine can enforce, and an auditor can cite. Scattered across prompts, credentials, roles, approval tools and budgets, it is none of these. The mandate proposes one object compiled into deterministic enforcement at the point of action. Its value is testable: fewer authority failures, fewer unnecessary approvals, and faster revocation, measured on benchmarks built for the purpose.

FAQ

How is a mandate different from an OAuth scope? A scope says which resources a token can access. A mandate also carries purpose, budget, autonomy per action class, verifier requirements, delegation rules and expiry. It compiles into scoped tokens rather than replacing them.

Can the agent change its own mandate? No. Changing a mandate, or widening delegation, is never delegable to an agent.

What happens to an action in progress when a mandate is revoked? New consequential actions are denied at the next online authorization check. An action already dispatched may have an unknown outcome and must be reconciled, not retried.

References

  1. Hardt, D. (2012). RFC 6749: The OAuth 2.0 Authorization Framework. https://www.rfc-editor.org/rfc/rfc6749
  2. Jones, M. et al. (2020). RFC 8693: OAuth 2.0 Token Exchange. https://www.rfc-editor.org/rfc/rfc8693
  3. Lodderstedt, T., Richer, J., Campbell, B. (2023). RFC 9396: OAuth 2.0 Rich Authorization Requests. https://www.rfc-editor.org/rfc/rfc9396
  4. Birgisson, A. et al. (2014). Macaroons: Cookies with Contextual Caveats for Decentralized Authorization in the Cloud. NDSS 2014. https://www.ndss-symposium.org/ndss2014/programme/macaroons-cookies-contextual-caveats-decentralized-authorization-cloud/
  5. SPIFFE project. Secure Production Identity Framework for Everyone. https://spiffe.io/docs/latest/spiffe-about/overview/
  6. Debenedetti, E. et al. (2024). AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. arXiv:2406.13352. https://arxiv.org/abs/2406.13352
  7. Levy, I. et al. (2024). ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents. arXiv:2410.06703. https://arxiv.org/abs/2410.06703
  8. Ruan, Y. et al. (2023). Identifying the Risks of LM Agents with an LM-Emulated Sandbox. arXiv:2309.15817. https://arxiv.org/abs/2309.15817
  9. Parasuraman, R., Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors 52(3):381–410. https://doi.org/10.1177/0018720810376055
  10. Ethen Research Lab (2026). AgentTrustBench system card: autonomy boundaries on one pinned build (run R08-20260921-01). Published system card, /resources/research/agent-trust-boundaries.

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Developers

    How Ethen Gateway Binds API Requests to Projects

    Every chat request carries a key, lands in exactly one project, and leaves a paper trail. Here is how the Gateway enforces that chain.

  • Security & Trust

    Binding Computer-Use Approvals to Specific Actions

    An approval that says "yes" to the wrong action is worse than no approval at all. Here is how Ethen's computer-use path ties each decision to one exact action.

  • Product

    Ethen Code: What We're Building Next

    Ethen Code is moving from assistive coding — explaining code, writing functions, fixing snippets — toward software tasks that end with evidence a person can check. The direction has seven parts: reproduce a problem before changing anything; plan a minimal, reviewable change before executing it; work in isolated environments with scoped permissions; treat tests, builds and review as evidence rather than as a finish line; put approvals in front of merges, deployments and other consequential steps; recover from failures in long-running work without repeating effects; and keep the same task model across Chat, the cloud workspace and Desktop. This article describes that direction. It is not a release schedule, and it makes no availability or date claims.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.