Technical Report · 2026-10-03 · Trust & Accountable AI Work
Work Receipts: A Verifiable Record for Autonomous AI Work
When an AI agent does work, five different parties want to know what happened. Each usually keeps its own record, and the records disagree. We propose one signed receipt per task that all five can trust.
Abstract
Autonomous AI agents now take consequential actions: they modify code, update records, send messages and move money. Each such unit of work is recorded several times over. There are logs for engineers, usage meters for billing, audit trails for compliance, labeled examples for evaluation, and datasets for training. These records are built separately, at different granularities, and they drift apart. This technical report proposes the Work Receipt, a per-task AI agent audit trail and accounting record that joins authority, actors, actions, observed effects, verification, outcome state, cost and reuse rights under one signed, hash-linked envelope. We specify the receipt's field families and the invariants it must satisfy. We describe how one receipt can serve billing, audit, evaluation, training and incident reconstruction, and we discuss tamper-evidence levels and privacy. We also state what a receipt cannot prove: a signature establishes that a record has not been altered, not that its claims are true or that every event was captured. The receipt is an Ethen architecture proposal. It is not yet implemented as described, and this report contains no measured results.
The problem: five records of one event
Consider an agent that resolves a billing dispute. It reads a ticket, queries a customer's invoices, drafts a refund, obtains approval for an amount above a threshold, issues the refund through a payment API, updates the ticket and notifies the customer. At least five systems record this work:
- Engineering telemetry records model calls, tool calls, latencies and errors.
- Metering records tokens and API usage for cost and billing.
- Audit records who approved what, under which policy.
- Evaluation records whether the dispute was correctly resolved, if anyone checks.
- Data pipelines record an example that may later be used for training, if rights permit.
Each record answers a narrow question. None can answer the question that matters when something goes wrong: who authorized this action, what exactly happened in the world, how do we know the outcome, what did it cost, and are we permitted to learn from it? Answering that requires joining records that were never designed to be joined. The join is usually done by hand, after an incident, from fragments that disagree.
An AI agent audit trail that cannot be reconciled with billing or evaluation is a weak audit trail. A billing record that cannot point to verification is a weak basis for charging. A training example that cannot point to its rights is a liability.
Proposal: one receipt per task
A Work Receipt is a single record emitted when a task reaches a terminal state, or a pending state that is explicitly declared. It is the authoritative join of everything needed to account for that task. Figure 1 shows its field families.
Figure 1. Anatomy of a Work Receipt. The receipt groups eleven field families under one signed envelope. Content is referenced by hash, never embedded. Field names are illustrative; this is a design proposal, not a published schema. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (Work Receipt); not implemented as described.
An illustrative schema:
WorkReceipt {
receipt_id, tenant_id, task_id, run_ids[], schema_version
spec_ref, spec_hash # what was asked; content by reference only
mandate_ref # authority: scope, budget, approvals, expiry
actors[] # agent principals, models@version, tools@version, humans
routing[] # decisions, candidate sets, selection probabilities
actions[] # proposed → admitted → (approved) → executed → reconciled
effects[] # observed state changes; external receipts/IDs
verification { # how the outcome is known
level: invariant | programmatic | expert | judge
verifier_id@version, verdict, confidence,
calibrated_fpr, calibrated_fnr, calibration_set_version
}
outcome_state # success | partial | failure | unknown |
# rejected | rolled_back | superseded
cost { model, tools, sandbox, infra, verification, human_review }
rights { purposes_allowed[], retention, residency, revocation_ref }
evidence_chain_head # hash of last per-event record
signatures[]
}Five design choices follow from the problem statement.
References, not content. The receipt carries content identifiers and hashes, never raw prompts, documents or outputs. Content stays in tenant-controlled stores under its own retention rules. This keeps receipts small and avoids building a hidden cross-tenant content lake. It also lets a receipt outlive the deletion of the content it describes: the receipt remains, recording that the work occurred, while the content is gone.
Authority is a first-class field. Every receipt points to the mandate under which the work ran: the human or organizational grant that bounded its scope, budget and required approvals. A receipt without authority cannot answer the audit question.
Actions have a lifecycle. Side effects are recorded through their states: proposed, admitted by policy, approved where required, executed, and reconciled against the system of record. An action whose execution status could not be confirmed is recorded as such. The note on unknown effects explains why "timed out" must never be recorded as "failed". Where a multi-step action cannot be rolled back, the receipt records the compensating actions taken instead, following the long-established pattern for long-lived transactions (Garcia-Molina & Salem).
Verification is explicit. The outcome names the verifier that produced it, the verifier's version, and that verifier's calibrated error rates as of the most recent calibration. Downstream consumers can then weight or exclude labels by trust tier. The measurement behind those error rates is the subject of Evaluating the Evaluators.
Rights travel with the record. The purposes for which the task's experience may be reused are attached at creation, not inferred later. See Rights as Infrastructure.
One record, five consumers
The value of the receipt comes less from any single field than from the fact that five consumers read the same record (Figure 2).
Figure 2. One record, five consumers. Billing, audit, evaluation, training and incident reconstruction usually maintain separate records that disagree. A single receipt gives them one source of truth, and the disagreements become detectable conformance failures rather than reconciliation projects. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen architecture proposal (Work Receipt); not implemented as described.
Billing. If work is priced per verified unit, the charge depends on verification.verdict and outcome_state, and the cost fields support cost-per-outcome accounting. The economic unit this enables is developed in Cost Per Verified Outcome.
Audit and compliance. The mandate reference, actor list, approval states and evidence chain answer the core audit questions. Regulatory record-keeping obligations for some AI systems, such as the logging requirements of the EU AI Act for high-risk systems, may be easier to evidence from a receipt than from ad hoc logs. Whether and how any specific obligation applies to a given deployment is a question for counsel, not for this report.
Evaluation. A receipt is a labeled trajectory whose label has known provenance; in the terms of From AI Traces to Verified Experience, it is the container for verified experience. Evaluation sets can be assembled from receipts at a stated verification level, with contamination controls applied at build time. Because cost and outcome sit in the same record, results can be reported with cost alongside accuracy, as recommended for agent evaluation (Kapoor et al.), and summarized in established reporting formats such as model cards (Mitchell et al.).
Training. A receipt carries exactly the metadata a training build needs to decide eligibility: rights, label tier, verifier error and lineage. The Outcome Warehouse is the analytical store over receipts. The rights-aware dataset compiler turns receipts into datasets.
Incident reconstruction. Because the receipt commits to the head of a hash-linked chain of per-event records, an investigator can reconstruct the exact sequence of decisions, approvals and effects and detect whether any record was altered.
This is also where we locate the defensibility of the design, to the extent there is any. Any one of these records is easy to build. Keeping five consumers in agreement on one signed object across many tasks is an engineering and organizational discipline. That discipline compounds over time.
Three record layers
Receipts are one layer of a three-layer evidence contract (Figure 3). Per-event records are emitted at each state transition: task created, routing decided, action proposed, approval decided, action executed, outcome observed. Each event carries actor, policy version, time and a hash link to its predecessor. The receipt is the per-task summary that commits to the chain head. An incident bundle collects receipts, policy snapshots, identity state and reconciled effects when something needs investigation.
Figure 3. From events to receipts to incident bundles. Three record layers. Per-event records are emitted at every state transition; the receipt summarizes a run and commits to its event chain; an incident bundle gathers receipts, policy snapshots and reconciled effects for investigation. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (three-record evidence contract).
Invariants
A receipt system is only useful if it is held to invariants that can be tested. We propose the following as conformance requirements:
- Uniqueness. Every task that reaches a terminal or declared-pending state has exactly one current receipt. Later receipts supersede earlier ones and never overwrite them.
- Cost reconciliation. The receipt's cost fields reconcile with the independent usage meter within a stated tolerance.
- Action completeness. Every consequential action event in the chain appears in the receipt's action list, and vice versa.
- Authority binding. Every executed action that required approval references an approval bound to that exact action, its parameters and the policy version in force.
- Rights completeness. Every receipt carries a rights record. A missing rights record means no reuse permitted, not unknown.
- Verification provenance. Every non-pending outcome names a verifier and its version. Model-judge verdicts are never the sole basis for a billed outcome.
These invariants are testable in a conformance suite that injects crashes before and after effects, duplicate deliveries, stale workers and revoked approvals, then checks that receipts remain correct.
Tamper evidence: what signatures can and cannot show
How strongly should receipts resist alteration? Transparency-log practice suggests a ladder (Figure 4), built from the Merkle-tree machinery of Certificate Transparency (RFC 6962; RFC 9162) and external timestamping (RFC 3161).
Figure 4. Tamper-evidence levels for agent evidence. A ladder adapted from transparency-log practice (RFC 6962 / RFC 9162). Higher levels make silent rewriting harder, but no level shows that a log is complete. Completeness needs independent emission and reconciliation. Evidence label: QUALITATIVE MATRIX. Source: Adapted from RFC 6962/9162 transparency-log design; Ethen internal synthesis.
An append-only hash chain (L1) detects edits to individual records. Merkle batching with signed checkpoints (L2) makes it hard for an operator with database access to rewrite history without detection, provided checkpoints are retained somewhere the operator does not control. Witness cosigning (L3) prevents an operator from showing different histories to different parties. External anchoring (L4) establishes when a checkpoint existed. Our working proposal is L1 from the start, L2 for enterprise deployments, and L3 as a research target with customer- or auditor-held witness keys. We have not implemented any of these levels as described.
Two limits should be stated plainly. A valid signature establishes the integrity and attribution of a record, not the truth of its contents. If an agent's runtime records that a refund succeeded when the payment provider never processed it, a perfectly signed receipt is still wrong. Truth requires independent observation: reconciliation against the system of record, and external receipts from the provider. And no level of the ladder proves completeness. A log can be immaculate and still omit events that were never emitted. Completeness requires independent emission paths and periodic reconciliation, for example checking that every irreversible action in a payment system's own records has a corresponding approved receipt.
Interoperability
Receipts should not require a proprietary ecosystem. Several open standards cover parts of the problem. W3C PROV-O provides a vocabulary for entities, activities and agents that can express receipt lineage. in-toto provides signed attestations about steps in a supply chain (Torres-Arias et al.). OpenTelemetry's generative-AI conventions standardize span attributes for model and tool calls. A reasonable design publishes the receipt schema openly and maps it to these standards. The derivation logic that produces receipts from a particular runtime can remain implementation detail.
Failure modes and open questions
- Overhead. Signing, hashing and reconciling every task adds latency and storage. Whether that cost is acceptable at high volume is an empirical question.
- Granularity mismatch. Billing may want per-task records while incident response wants per-step detail. The three-layer design addresses this in principle; in practice consumers may need incompatible aggregations.
- Delayed outcomes. Many outcomes are not known when a task ends. Receipts must support pending states and superseding records without breaking consumers that expect finality.
- Verifier conflict of interest. If an operator bills on its own verifiers, it grades its own work. Mitigations include customer-pinned verifier versions, a dispute channel that uses receipts as evidence, periodic third-party calibration audits, and a preference for deterministic checks for billed outcomes.
- Privacy leakage through metadata. Even without content, receipts reveal actors, timing, tool sequences and resource identifiers. Metadata can identify customers and individuals, so receipts must be treated as tenant-private by default.
Limitations
This report proposes a design. It has not been implemented end to end, and no measurement supports its claimed benefits for audit speed, billing accuracy or label quality. The field list is illustrative and will change. The legal value of receipts for any regulatory obligation is undetermined and requires counsel review. The proposal assumes that the runtime can observe and reconcile effects, which is not true of every external system.
Conclusion
Agents that act need records that can be trusted by everyone who depends on them. The Work Receipt is a proposal to stop maintaining five partial records of the same work and maintain one: authority, actions, effects, verification, cost and rights, signed and hash-linked, with content held by reference. Its strongest claim is modest and testable: a single receipt, held to explicit invariants, makes disagreements between billing, audit, evaluation and training visible as conformance failures instead of invisible as drift.
FAQ
Is a Work Receipt the same as an audit log? No. An audit log records events. A receipt is a per-task summary that joins authority, effects, verification, cost and rights, and commits to the underlying event log by hash.
Does a receipt contain the customer's data? No. It contains references and hashes. Content stays in tenant-controlled storage under its own retention and deletion rules.
Does a signed receipt prove the agent did the right thing? No. A signature shows that the record has not been altered since it was signed. Whether the outcome was correct depends on the verifier and on independent reconciliation.
Will the schema be open? We think the schema should be openly documented so that other runtimes can emit compatible receipts. That decision has not been made.
Related research
- Mandates: Compiling Human Intent Into Bounded Agent Authority — the authority a receipt points to.
- From AI Traces to Verified Experience — receipts carry verified experience.
- The Outcome Warehouse: Turning Completed AI Work Into Research Assets — receipts feed the outcome warehouse.
- Rights as Infrastructure: Building AI Datasets That Know How They May Be Used — rights fields in the receipt.
- Cost Per Verified Outcome: A Better Economic Unit for Agentic AI — receipts as the CPVO numerator and denominator.
- Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry — unknown effects recorded honestly.
- Evaluating the Evaluators: Reward Integrity for AI Agents — verification fields and verifier independence.
References
- Laurie, B., Langley, A., Kasper, E. (2013). RFC 6962: Certificate Transparency. https://www.rfc-editor.org/rfc/rfc6962
- Laurie, B., Messeri, E., Stradling, R. (2021). RFC 9162: Certificate Transparency Version 2.0. https://www.rfc-editor.org/rfc/rfc9162
- Adams, C. et al. (2001). RFC 3161: Internet X.509 Public Key Infrastructure Time-Stamp Protocol. https://www.rfc-editor.org/rfc/rfc3161
- W3C (2013). PROV-O: The PROV Ontology. https://www.w3.org/TR/prov-o/
- Torres-Arias, S. et al. (2019). in-toto: Providing farm-to-table guarantees for bits and bytes. USENIX Security 2019. https://www.usenix.org/conference/usenixsecurity19/presentation/torres-arias
- OpenTelemetry. Semantic conventions for generative AI systems. https://opentelemetry.io/docs/specs/semconv/gen-ai/
- Regulation (EU) 2024/1689 (Artificial Intelligence Act). Official Journal of the EU, 12 July 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
- NIST (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. https://doi.org/10.6028/NIST.AI.100-1
- Garcia-Molina, H., Salem, K. (1987). Sagas. Proc. ACM SIGMOD. https://doi.org/10.1145/38713.38742
- Kapoor, S. et al. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
- Mitchell, M. et al. (2018). Model Cards for Model Reporting. arXiv:1810.03993. https://arxiv.org/abs/1810.03993
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Evaluating the Evaluators: Reward Integrity for AI Agents
A methods paper on reward integrity for AI agents: verifier false accepts and rejects, abstention, grader drift, expert disagreement and reward hacking.
- Counterfactual Replay for AI Agents
A research proposal for counterfactual evaluation of AI agents: replaying completed tasks under alternative models, tools, context and recovery strategies.
- Why Better Foundation Models May Make Evaluation More Valuable, Not Less
A position paper stress-testing AI evaluation against 10× better models and 10× cheaper inference, and arguing that verification and assurance gain value.
Explained on the Ethen Blog
- Making Mission Completion Depend on Evidence
In the Stage-0 mission system, running the work and proving the work are different transitions — and only independent evidence unlocks the second.
- What Makes an AI Agent Job Verifiable
An agent job is verifiable when someone other than the agent can confirm what was supposed to happen, what actually happened, and what remains unknown. This guide turns that idea into a checklist you can apply to any system.
- How to Explore Ethen Research Lab: Programs, Evidence Labels and Reading Paths
The fastest way to read Ethen Research Lab well is to check two labels before reading anything else: the publication type (position paper, research note, proposal, technical report, methods paper, protocol, benchmark design, survey or system card) and the evidence status (measured result, synthesis, proposal, protocol, or external survey). Together they tell you what kind of claim the paper can make. Then filter the archive by research program to find papers on your topic, and use a reading path to follow a question from concept to benchmark to experiment. This guide explains each label, the programs, how related papers fit together, and where to start for your role.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.