Research Note · 2026-10-03 · Trust & Accountable AI Work
Mandates: Compiling Human Intent Into Bounded Agent Authority
An agent's authority today is scattered across prompts, API keys, role settings, approval rules and budget caps. We propose one object, the mandate, from which all of them are compiled.
Abstract
When a person asks an AI agent to "clear this week's billing backlog, but check with me before refunding more than $200", they express a purpose, a scope, a budget, a risk threshold and an approval rule in one sentence. Current systems implement that sentence in pieces: a system prompt, an API credential, a role in an identity provider, an approval workflow and a spending cap, each configured separately and none aware of the others. This note proposes the mandate: a single, explicit object that records what a human principal has authorized an agent to do. It is then compiled into the enforcement artifacts that operate where the agent acts: attenuated capability tokens, constraints on allowed configurations, budget meters and approval routing. We describe the mandate's fields and its autonomy levels per action class. We describe delegation with attenuation, expiry and revocation, and the conditions under which AI agent authorization can be tested rather than assumed. We cite one published, narrowly scoped Ethen measurement of boundary enforcement. Everything else here is an architecture proposal.
The problem: authority with no single source
An agent that acts on behalf of a person or organization needs authority. In most deployments that authority is assembled from five mechanisms built independently:
- Instructions in the system prompt ("never refund more than $200").
- Credentials that grant access to tools and data, often broader than the task needs.
- Roles in an identity system, designed for humans.
- Approval workflows that pause for human confirmation, configured per tool.
- Budgets for model spend, often enforced at the account level rather than the task.
Each mechanism answers a different question, and none answers the one that matters: what exactly did this person authorize this agent to do, for how long, and under what conditions? The gaps between mechanisms are where failures occur. An instruction can be overridden by injected content. A credential can permit what the instruction forbids. An approval can be granted for one action and then applied to a modified version. A budget can be exhausted by a different task in the same account.
The first of these is the most serious. Prompt injection, in which untrusted content read by an agent alters its behavior, remains an open problem, and benchmarks such as AgentDojo exist precisely because defenses are incomplete (Debenedetti et al.). Policy-compliance benchmarks for web agents show that task success and policy adherence can diverge (Levy et al.). Emulated sandboxes have been used to surface risky agent actions before they reach real tools (Ruan et al.), which is useful for testing mandates but does not replace enforcing them. An authority model that lives in the prompt is an authority model that injected text can reach.
Proposal: the mandate
A mandate is a structured grant of authority from a principal to an agent. In the illustrative schema below, the principal is a human or an organizational role.
Mandate {
mandate_id, principal, purpose # who grants, and why
scope { resources, systems, data_purposes }
budget { money, tokens, wall_time, human_review_minutes }
autonomy[action_class] ∈ { auto, notify, approve, forbidden }
verifier_requirements # what evidence closes the task
expiry, revocation_epoch
delegation { max_depth ≤ 3, allowed_child_scopes }
}The mandate is not enforced by the model. It is compiled (Figure 1) into four artifacts, each enforced by deterministic machinery at the point of action:
Figure 1. A mandate compiles into four enforcement artifacts. The human-facing mandate is one object. It compiles into enforcement artifacts that operate where the agent acts: attenuated capability tokens at tool boundaries, constraints on which execution configurations are allowed, budget meters, and approval routing. None of these is enforced by the model's prompt. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (Mandate); not implemented as described.
- Capability tokens. Short-lived, audience-bound credentials carrying only the permissions the mandate allows for this task. Attenuation, the property that a holder can narrow but never widen a credential, has a long history in authorization design, including macaroons with contextual caveats (Birgisson et al.).
- Configuration constraints. Limits on which models, providers, tools and sandboxes the agent's decision layer may choose. A mandate that forbids sending data outside a region forbids routing to a provider in another region. Action authority of this kind is distinct from data-use rights, which govern what may later be learned from the work; see Rights as Infrastructure. See Faros.
- Budget meters. Reservations against money, tokens, time and human-review minutes, checked before each consequential step rather than reconciled afterward.
- Approval routing. Rules that determine which action classes pause for whom, and with what second factor.
This framing has a practical consequence: the user-facing interaction ("give the agent a mandate"), the security model and the cost model become the same object. A person can read what they authorized. An enforcement layer can check it. A Work Receipt can cite it.
Relation to existing standards
The mandate is not a new protocol. It sits above existing authorization standards and compiles into them. OAuth 2.0 defines delegated access with scopes (RFC 6749). OAuth Token Exchange allows a token to be exchanged for another with different, typically narrower, properties, which suits multi-hop delegation (RFC 8693). OAuth Rich Authorization Requests express fine-grained, structured authorization details beyond flat scope strings (RFC 9396), and are perhaps the closest existing standard to the mandate's scope field. Workload identity frameworks such as SPIFFE provide service identities for the software that executes agents. A mandate system should map onto these standards rather than replace them. What the standards do not provide is the agent-level object: purpose, autonomy per action class, verifier requirements and a delegation policy, bound together and compiled consistently.
Autonomy levels per action class
Prompting a human for every step defeats the purpose of delegation and invites approval fatigue: people asked to confirm too often stop reading what they confirm. Human-factors research on automation describes related complacency and bias effects (Parasuraman & Manzey). Granting full autonomy is unacceptable for consequential actions. The mandate resolves this tension by setting autonomy per action class rather than per agent (Figure 2).
Figure 2. Autonomy levels by action class: a proposed default. Autonomy is set per action class, not per agent. Irreversible or authority-changing actions are never automatic; reversible, in-scope writes can run under standing bounded authority. These defaults are a proposal from Ethen's internal architecture work, to be tested in usability and safety studies. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal; proposed defaults, untested.
The proposed defaults reflect one principle: strict where loss is irreversible, permissive where work is reversible and in scope. Money movement always requires explicit approval, with a stronger second factor for high-risk amounts. Permission, policy and key changes are never delegated to an agent at all. Reversible writes within an explicitly bounded scope, such as drafts, internal notes and bounded record updates, can proceed under standing authority with audit. These are defaults to be studied, not established findings. The relationship between autonomy settings, approval experience and outcomes is an open research question.
Approvals must bind to exact actions
An approval is meaningful only if it applies to the action that is actually executed. Mandate-based approval therefore binds to the action's exact parameters, the principal, the environment, the policy version in force, the resource scope, an expiry, and a hash of the material content. If the agent or the environment changes anything material after approval, the approval no longer applies and must be re-requested. This closes a common gap in which a harmless plan is approved and a different effect is executed after a model edit, a retry, or a change in the target's state. The interaction between approval binding and ambiguous outcomes is discussed in Unknown Effects in Autonomous AI Systems.
Delegation narrows authority
Agents increasingly delegate to subagents. Each delegation hop is an opportunity for authority to widen accidentally, through a confused deputy that uses its own broader credentials on a caller's behalf. The mandate's delegation rules (Figure 3) are designed to make widening structurally impossible.
Figure 3. Delegation narrows authority at every hop. Each subagent receives a distinct principal and a grant that is the intersection of its parent's grant and its own declared need. Depth is capped. A fixed list of permissions can never be delegated, whatever the mandate says. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (delegation attenuation; never-delegable list).
Every subagent receives a distinct principal and a grant equal to the intersection of its parent's grant and its own declared need. Delegation depth is capped; the internal proposal is three. A fixed list of permissions is never delegable whatever the mandate says: identity administration, policy writes, key export, impersonation, cross-tenant access, granting approvals, overriding data residency, mutating evidence, and widening delegation. No agent may approve its own privilege expansion.
Expiry and revocation
Authority that cannot be withdrawn is not bounded. Mandates expire. They can also be revoked at any time by the principal or an administrator. Token lifetimes alone cannot deliver fast revocation: a fifteen-minute token remains valid for up to fifteen minutes after the mandate behind it is withdrawn. The design therefore pairs short-lived tokens with an online authorization check immediately before each consequential action, so that revocation takes effect at the next action rather than at the next token refresh.
Ethen's internal architecture work proposes engineering targets for study [PROPOSED TARGET]: stopping an active agent within about five seconds of a kill request; denying new consequential actions within thirty seconds at the 99th percentile after revocation; token lifetimes of fifteen minutes or less; and authorization-epoch rotation within five minutes. These are targets for measurement under a defined threat model, not implemented service levels. Actions already in flight when revocation arrives raise the same ambiguity as any interrupted effect.
What has been measured
Ethen has published one narrowly scoped measurement relevant to mandate enforcement: the AgentTrustBench system card for run R08. On one pinned build, 197 enumerated conditions covering allow and deny decisions, stop conditions, spend limits, approval binding and recovery were each run once and then replayed in a different order. [MEASURED ETHEN RESULT] 194 of 197 conditions matched both the pre-registered prediction and an independent oracle on both passes. The three non-matching conditions are described in the card. As the card itself states, the result applies only to that build and that enumerated set. It supports no population claim, no product ranking and no claim about live traffic, and no statistical test was applied. It is cited here as an example of how mandate-style boundaries can be tested, not as evidence that the full design in this note works. The benchmark design that would test mandates broadly is VerifiedWork Control.
Mandates and commitments
A mandate states what an agent may do. It does not state what the agent must accomplish, or what remains unfinished. That is the role of a commitment graph: the record of obligations, preconditions and evidence that determines when work is actually complete. The two are complementary. Mandate violations are failures of authority; commitment violations are failures of completion. A system needs both to say, at the end of a task, "this was allowed, this was done, and this is how we know."
Failure modes and open questions
- Compiling natural language is itself error-prone. A model that translates "clear the billing backlog" into a mandate can misread scope. The compiled mandate must be shown to the principal in a reviewable form before it takes effect.
- Over-broad mandates. Principals may grant wide mandates for convenience, recreating the problem the mandate was meant to address. Interfaces that make scope legible, and defaults that start narrow, are design obligations. Neither is solved.
- Action-class taxonomy drift. New tools introduce actions that do not fit existing classes. An unclassified action should default to the most restrictive class.
- Approval fatigue. Even per-class autonomy may produce too many approvals in some workflows. Whether mandates reduce approvals per task at constant safety is an empirical question.
- Usability. Users may find mandates harder to reason about than per-action prompts. That would argue for changing the design, not for weakening enforcement.
Limitations
This note describes an architecture that has not been implemented as described. The only Ethen measurement cited covers one pinned build and one enumerated condition set; it does not validate the design here. The proposed autonomy defaults and revocation targets are untested. The mandate addresses authority; it does not, by itself, defend against every way an agent can misuse authority it legitimately holds.
Conclusion
Agent authority should be something a person can read, a machine can enforce, and an auditor can cite. Scattered across prompts, credentials, roles, approval tools and budgets, it is none of these. The mandate proposes one object compiled into deterministic enforcement at the point of action. Its value is testable: fewer authority failures, fewer unnecessary approvals, and faster revocation, measured on benchmarks built for the purpose.
FAQ
How is a mandate different from an OAuth scope? A scope says which resources a token can access. A mandate also carries purpose, budget, autonomy per action class, verifier requirements, delegation rules and expiry. It compiles into scoped tokens rather than replacing them.
Can the agent change its own mandate? No. Changing a mandate, or widening delegation, is never delegable to an agent.
What happens to an action in progress when a mandate is revoked? New consequential actions are denied at the next online authorization check. An action already dispatched may have an unknown outcome and must be reconciled, not retried.
Related research
- Work Receipts: A Verifiable Record for Autonomous AI Work — receipts reference the mandate.
- Commitment Graphs: Why AI Agents Need to Know What Is Still Unfinished — mandates authorize, commitments obligate.
- VerifiedWork Control: Evaluating Delegation, Approval, Revocation, and Agent Authority — benchmark track that tests mandate enforcement.
- Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry — approval binding under unknown effects.
- Faros: Researching How Intelligence Should Choose Intelligence — mandates constrain Faros choices.
- Rights as Infrastructure: Building AI Datasets That Know How They May Be Used — data-purpose rights alongside action authority.
References
- Hardt, D. (2012). RFC 6749: The OAuth 2.0 Authorization Framework. https://www.rfc-editor.org/rfc/rfc6749
- Jones, M. et al. (2020). RFC 8693: OAuth 2.0 Token Exchange. https://www.rfc-editor.org/rfc/rfc8693
- Lodderstedt, T., Richer, J., Campbell, B. (2023). RFC 9396: OAuth 2.0 Rich Authorization Requests. https://www.rfc-editor.org/rfc/rfc9396
- Birgisson, A. et al. (2014). Macaroons: Cookies with Contextual Caveats for Decentralized Authorization in the Cloud. NDSS 2014. https://www.ndss-symposium.org/ndss2014/programme/macaroons-cookies-contextual-caveats-decentralized-authorization-cloud/
- SPIFFE project. Secure Production Identity Framework for Everyone. https://spiffe.io/docs/latest/spiffe-about/overview/
- Debenedetti, E. et al. (2024). AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. arXiv:2406.13352. https://arxiv.org/abs/2406.13352
- Levy, I. et al. (2024). ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents. arXiv:2410.06703. https://arxiv.org/abs/2410.06703
- Ruan, Y. et al. (2023). Identifying the Risks of LM Agents with an LM-Emulated Sandbox. arXiv:2309.15817. https://arxiv.org/abs/2309.15817
- Parasuraman, R., Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors 52(3):381–410. https://doi.org/10.1177/0018720810376055
- Ethen Research Lab (2026). AgentTrustBench system card: autonomy boundaries on one pinned build (run R08-20260921-01). Published system card, /resources/research/agent-trust-boundaries.
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- AgentTrustBench system card: autonomy boundaries on one pinned build
System card for one offline enumeration of a pinned Ethen build: 197 conditions, 194 passed. Not a paper and not a security disclosure.
- Work Receipts: A Verifiable Record for Autonomous AI Work
A technical report proposing the Work Receipt: one signed record of authority, actions, effects, verification, cost and rights for every unit of autonomous AI work.
- Rights as Infrastructure: Building AI Datasets That Know How They May Be Used
A research note on AI training data rights as infrastructure: purpose grants, consent, residency, expiry, revocation, descendants and training eligibility.
Explained on the Ethen Blog
- How Ethen Gateway Binds API Requests to Projects
Every chat request carries a key, lands in exactly one project, and leaves a paper trail. Here is how the Gateway enforces that chain.
- Binding Computer-Use Approvals to Specific Actions
An approval that says "yes" to the wrong action is worse than no approval at all. Here is how Ethen's computer-use path ties each decision to one exact action.
- Ethen Code: What We're Building Next
Ethen Code is moving from assistive coding — explaining code, writing functions, fixing snippets — toward software tasks that end with evidence a person can check. The direction has seven parts: reproduce a problem before changing anything; plan a minimal, reviewable change before executing it; work in isolated environments with scoped permissions; treat tests, builds and review as evidence rather than as a finish line; put approvals in front of merges, deployments and other consequential steps; recover from failures in long-running work without repeating effects; and keep the same task model across Chat, the cloud workspace and Desktop. This article describes that direction. It is not a release schedule, and it makes no availability or date claims.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.