Benchmark Design · 2026-10-03 · Evaluation & Verification
VerifiedWork Context: Measuring What AI Agents Must Remember
Long-context benchmarks ask whether a model can find something in its input. Agents face a harder test: whether they still honor what they were told hours and several compactions ago, and whether they have stopped using what they were told to forget.
Abstract
Long-running agents must remember some things and forget others. They must keep honoring a constraint stated early in a session, fulfill obligations accumulated along the way, and quote evidence accurately. They must also stop using content whose permission was revoked, facts that were superseded and sources that were deleted. Existing long-context and memory benchmarks measure parts of this: retrieval from long inputs, multi-session recall, knowledge updates, selective forgetting. VerifiedWork Context is the context track of Ethen VerifiedWork, an agent memory benchmark built around what an agent must remember to act correctly. Tasks plant six kinds of probe at controlled points, namely session constraints, obligations, permission revocations, superseded facts, deleted sources and required quotations. Context compaction is forced at controlled moments, and every task ends in a behavioral test. The primary metric is whether the agent's actions comply with everything it should still remember and nothing it should have dropped. Retention questions are reported separately as a diagnostic, because remembering a constraint and following it are different things. Token cost is reported alongside. The track is a proposed design and has not been run.
What existing benchmarks measure
Evaluation of long context and memory has become more demanding. Needle-in-a-haystack retrieval has given way to benchmarks with multiple needles, multi-hop tracing and aggregation. On one such benchmark, almost all models showed large drops as context length grew, despite near-perfect scores on simple retrieval (Hsieh et al.). When needles share little vocabulary with questions, performance degrades sharply with length (Modarressi et al.). Models also tend to miss information placed in the middle of long inputs (Liu et al.).
Memory benchmarks for assistants and agents test persistence across sessions. LoCoMo provides very long conversations grounded in persona and event graphs (Maharana et al.). LongMemEval evaluates five long-term memory abilities, including knowledge updates and abstention, and reports that commercial assistants and long-context models showed a 30% accuracy drop on memorizing information across sustained interactions [EXTERNAL PRIMARY-SOURCE RESULT] (Wu et al.). MemoryAgentBench identifies four competencies for memory agents: accurate retrieval, test-time learning, long-range understanding and selective forgetting. It notes that earlier benchmarks did not cover all four (Hu et al.).
Closest to our concern, an evaluation of context compaction found that user-issued session constraints were retained only 17% of the time on average after compaction, across chat, agentic and research settings (Wang et al., Lost in Compaction).
Two gaps remain. Most benchmarks ask questions about remembered content rather than testing whether remembered content governs actions. And few test whether an agent stops using content it is no longer permitted to use. That second test is as much a security property as a memory property.
Six kinds of probe
Tasks in this track plant six kinds of probe (Figure 1).
Figure 1. Probe types and what counts as correct. Six probe types cover what agents must remember and what they must stop using. Each has a behavioral test; retention questions are asked separately as a diagnostic. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen benchmark design (proposed).
Session constraints are instructions meant to bind the rest of a session, such as "do not delete anything until I confirm". The test presents a situation where the constraint applies and checks the action.
Obligations are requirements added during the task, such as "also include last quarter's comparison". The test checks that they are fulfilled, or honestly reported as pending. This connects directly to the commitment graph.
Permission revocations withdraw the agent's access to a document or system partway through. The test checks that revoked content neither appears in nor influences later output. This is the property that the rights infrastructure is meant to enforce at the system level. The track tests whether it holds at the level of agent behavior.
Superseded facts are corrected during the task, such as a price that turns out to be wrong. The test checks that only the corrected value is used.
Deleted sources are removed during the task. The test checks that they are not cited afterward and that claims depending on them are flagged.
Evidence quotations are passages the final answer must quote verbatim. The test checks fidelity and attribution, because summaries routinely paraphrase exactly the wording that matters.
Anatomy of a task
Figure 2 shows a representative task timeline.
Figure 2. Anatomy of a context task. Events are planted at controlled points in a long task: a constraint stated early, a forced compaction, a revoked permission, a superseded fact and a deleted source. The test comes at the end, and it is behavioral: does the agent act in accordance with everything it should still remember, and nothing it should not? Evidence label: EXPERIMENT DESIGN. Source: Ethen benchmark design (proposed).
A task begins with a constraint stated early, followed by enough work to fill a substantial share of the context budget. Compaction is forced at a controlled point, so that every system under test compacts at the same moment rather than when its own heuristics decide. Revocations, supersessions and deletions occur at controlled later points. The behavioral test comes at the end. Probe phrasing and position vary systematically, because constraint retention has been shown to depend on phrasing and injection location (Wang et al.). Planting the same probe in different places and forms separates robust retention from luck.
Tasks also vary in session structure. Some run in one long session. Others span several sessions with memory carried between them, matching how assistants with persistent memory operate.
Retention versus use
The track distinguishes two measurements (Figure 3).
Figure 3. Retention is not use. A probe question shows whether information is still present; a behavioral test shows whether it governs action. The two can disagree in both directions, so the track reports them separately and treats the behavioral result as primary. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen benchmark design (proposed).
A retention probe asks the agent about the planted information, for example "what did the user ask you not to do?". A behavioral test presents a situation in which the information should govern action, without mentioning it. The two can disagree. An agent may recall the constraint when asked and still violate it when acting. It may also act correctly without being able to state the constraint, perhaps because the behavior was already its default. Behavioral compliance is the primary metric because it is what matters in deployment. Retention is a diagnostic that helps explain failures.
Metrics
For each system and condition the track reports:
- Behavioral compliance per probe type: the share of tests in which the agent's actions honored what it should remember and avoided what it should not.
- Revoked-content exposure: any appearance or measurable influence of revoked content after revocation. This is reported as a count with examples, because a single exposure may be unacceptable.
- Stale-fact use: use of superseded values.
- Evidence fidelity: exact-match rate of required quotations, with attribution.
- Verified task success, so that compliance is not achieved by refusing to do the work.
- Tokens and latency, so that compliance can be weighed against cost.
- Retention, reported separately as a diagnostic.
Scoring rules and statistics
Each probe in each task is scored independently, so one task yields several compliance observations. Two rules keep the headline honest. First, revoked-content exposure is critical: in task families that model regulated or confidential data, a single exposure fails the task regardless of everything else. Exposure counts are always reported as numbers of events, never only as rates. Second, compliance cannot be bought with refusal: an agent that avoids violating a constraint by not doing the task scores on compliance but fails on verified task success. Both are shown side by side.
Exposures and constraint violations should be rare, and rare events need large samples. If no exposures are observed in n independent tests, the upper 95% bound on the exposure rate is roughly 3/n (Hanley & Lippman-Hand). Reports therefore state the number of tests behind every "no exposures observed" claim and give the corresponding bound. Probes are clustered within tasks, so intervals are computed by resampling whole tasks.
Arms and baselines
The track evaluates any context strategy a system uses: full context where it fits, fixed summarization, retrieval, memory systems that page information in and out of context (Packer et al.) or track when facts were valid (Rasmussen et al.), or structured approaches such as the evidence-preserving context compiler. Each release reports baseline arms: full context (for tasks that fit), naive summarization at the forced compaction point, and retrieval-only over the task history. Comparisons are made at matched token budgets. A strategy that achieves high compliance by using far more tokens has not shown the same thing as one that achieves it within budget.
Permissions are part of memory
Treating permission revocation as a memory probe is a deliberate design choice. In agent systems, content enters context through retrieval, tool outputs, summaries and caches. When access is revoked, every one of those paths may still hold a copy. A system can enforce revocation at the retrieval layer and still leak revoked content through a summary produced before the revocation. Testing revocation as a behavioral property, rather than as an access-control check, catches leaks through every path at once. The track cannot detect every form of inference leakage. An agent may infer restricted facts from permitted ones, and the track reports that limitation explicitly.
An illustrative task
[ILLUSTRATIVE EXAMPLE — a design sketch, not a run.] An agent prepares a vendor-renewal recommendation over a long session. Early on, the user says: "Don't contact the vendor until I've approved the shortlist." The agent reads contracts, usage reports and pricing documents, and the context is compacted. The user's access to one pricing document is then revoked because it belongs to another department. A unit price in a usage report is corrected. A draft contract is deleted from the shared drive. At the end, the agent is asked to "finalize the recommendation and set up the next steps". A compliant agent drafts but does not send a vendor email, pending approval. It uses no figures from the revoked pricing document and only the corrected unit price. It does not cite the deleted draft, and it quotes the renewal clause from the remaining contract verbatim. Each of those is scored independently.
Relation to the rest of VerifiedWork
This track uses the environments and reporting standard of Ethen VerifiedWork. The experiment that estimates how much context different tasks actually need, across budgets, is How Should We Measure How Much Context an AI Agent Actually Needs?.
What this track cannot prove
A good score shows that an agent remembered and forgot correctly for the probe types and positions tested. It does not show that the agent never leaks restricted information through inference, that it handles every phrasing of a constraint, or that it retains information across arbitrarily many compactions. Chained degradation across many compaction cycles is under-studied, and the initial design includes only a small number of cycles.
Limitations
The track has not been built or run. Planted probes are artificial: real constraints are often implicit or ambiguous. Behavioral tests require situations that trigger the constraint without mentioning it, which takes careful task design and expert review. Measuring "influence" of revoked content, as distinct from its literal appearance, is difficult and will rely on paired comparisons with and without the content. Forcing compaction at fixed points makes systems comparable but may disadvantage systems whose own compaction timing is part of their design. Results are reported both with forced and with native compaction for that reason.
Conclusion
The memory that matters for agents is not the ability to find a fact in a long input. It is the ability to keep acting on the right constraints and to stop acting on what has been withdrawn, after the context has been compressed and the session has moved on. VerifiedWork Context tests that ability behaviorally, probe by probe, alongside the cost of achieving it.
FAQ
How is this different from needle-in-a-haystack tests? Needle tests ask the model to find information. This track tests whether remembered constraints govern later actions, and whether revoked or superseded information stops being used.
Why is permission revocation treated as a memory test? Because revoked content can persist in summaries, caches and earlier tool outputs. Testing behavior after revocation catches leaks through every path.
Does the track force compaction? Yes, at controlled points for comparability, with native-compaction results reported separately.
Related research
- Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action — umbrella benchmark.
- Evidence-Preserving Context: Compressing Agent Memory Without Losing Obligations — context compiler.
- How Should We Measure How Much Context an AI Agent Actually Needs? — context-need protocol.
- Commitment Graphs: Why AI Agents Need to Know What Is Still Unfinished — commitments as memory.
- Rights as Infrastructure: Building AI Datasets That Know How They May Be Used — revocation semantics.
References
- Wang, Z. et al. (2026). Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction. arXiv:2608.11242. https://arxiv.org/abs/2608.11242
- Hsieh, C.-P. et al. (2024). RULER: What's the Real Context Size of Your Long-Context Language Models? arXiv:2404.06654. https://arxiv.org/abs/2404.06654
- Modarressi, A. et al. (2025). NoLiMa: Long-Context Evaluation Beyond Literal Matching. arXiv:2502.05167. https://arxiv.org/abs/2502.05167
- Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172. https://arxiv.org/abs/2307.03172
- Maharana, A. et al. (2024). Evaluating Very Long-Term Conversational Memory of LLM Agents. arXiv:2402.17753. https://arxiv.org/abs/2402.17753
- Wu, D. et al. (2024). LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. arXiv:2410.10813. https://arxiv.org/abs/2410.10813
- Hu, Y. et al. (2025). Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions. arXiv:2507.05257. https://arxiv.org/abs/2507.05257
- Packer, C. et al. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560. https://arxiv.org/abs/2310.08560
- Rasmussen, P. et al. (2025). Zep: A Temporal Knowledge Graph Architecture for Agent Memory. arXiv:2501.13956. https://arxiv.org/abs/2501.13956
- Hanley, J. A., Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? JAMA 249(13):1743–1745. https://doi.org/10.1001/jama.1983.03330370053031
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Work Receipts: A Verifiable Record for Autonomous AI Work
A technical report proposing the Work Receipt: one signed record of authority, actions, effects, verification, cost and rights for every unit of autonomous AI work.
- Evaluating the Evaluators: Reward Integrity for AI Agents
A methods paper on reward integrity for AI agents: verifier false accepts and rejects, abstention, grader drift, expert disagreement and reward hacking.
- Counterfactual Replay for AI Agents
A research proposal for counterfactual evaluation of AI agents: replaying completed tasks under alternative models, tools, context and recovery strategies.
Explained on the Ethen Blog
- How We’re Rethinking AI Memory Across Ethen
AI memory should not be one opaque pile of things an assistant decided to remember. Ethen's direction for memory rests on five ideas. Different kinds of memory get different rules: your preferences, facts about you, a project's context, your organization's knowledge, learned procedures and commitments you made each have their own scope, lifetime and controls. Memory is not the record of what happened: the authoritative record of a task's actions and outcomes is kept separately and never replaced by a summary. Memory carries its source and its validity, so it can be checked, superseded and explained. Permissions travel with memory, so a summary of a restricted document stays restricted and revoking access revokes what was derived from it. And you can see, correct, export and delete what Ethen remembers, and see when a memory influenced what Ethen did. This article explains each idea and what it means across Ethen's apps. It describes direction, not shipped architecture.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.