Skip to content

EthenEthenEthen

Research Proposal · 2026-10-03 · Context / Skills / Transfer

Evidence-Preserving Context: Compressing Agent Memory Without Losing Obligations

Publication type
Research Proposal
Research program
Context / Skills / Transfer
Published
Authors
Ethen Research Lab
Reading time
12 min read

When an agent's context fills up, something has to go. What goes today is often the one instruction that mattered most. We propose deciding what is compressible before deciding how to compress it.

Cover image for "Evidence-Preserving Context: Compressing Agent Memory Without Losing Obligations". Decorative abstract motif; contains no data.

Abstract

Long-running agents accumulate more context than models can use well or afford. Context compaction, summarizing or dropping older material to stay within a budget, is therefore routine. It is also lossy in a dangerous way. A recent evaluation found that current compaction methods retained only 17% of user-issued session constraints, such as "do not delete any emails until I confirm", on average, and that most performed worse than running the same task without compaction. A constraint-aware extractor raised retention above 90%. We propose an evidence-preserving context compiler that generalizes this finding. Obligations, authority and permissions, deadlines, prerequisites for irreversible actions, approvals, active evidence and policy state travel in a protected structured channel that is never summarized. Everything else is compiled under a token budget into full, summarized, glimpsed or excluded form. Rights propagate: a summary inherits the most restrictive permissions of its sources, and revoking a source invalidates what was derived from it. We describe the design and its relation to long-context and memory research, and we propose an experiment comparing five strategies at matched budgets. The proposal is untested.

The problem: compaction forgets what matters

Context windows have grown, but the ability to use long contexts has not kept pace. Models can miss information placed in the middle of long inputs (Liu et al.). On a synthetic benchmark that goes beyond simple needle retrieval, almost all of 17 models evaluated showed large performance drops as context length increased, and only about half of those claiming 32,000-token contexts maintained satisfactory performance at that length [EXTERNAL PRIMARY-SOURCE RESULT] (Hsieh et al., RULER). When the needle and the question share little vocabulary, so that the model must infer an association rather than match words, 11 of 13 models evaluated fell below half of their short-context performance at 32,000 tokens [EXTERNAL PRIMARY-SOURCE RESULT] (Modarressi et al., NoLiMa). Long context is expensive as well as unreliable: cost and latency grow with every token.

The common response is compaction: when context approaches a limit, summarize older turns and tool outputs and continue. Summarization keeps the gist and drops specifics, and the specifics are often what an agent most needs. The clearest evidence comes from an evaluation of session constraints: instructions a user gives once that are meant to bind the rest of the session. Across multi-turn chat, agentic trajectories and long-horizon research, current compactors retained only 17% of injected session constraints on average, and most compacted runs performed worse than the same task without compaction [EXTERNAL PRIMARY-SOURCE RESULT] (Wang et al., Lost in Compaction). Retention varied with the compactor, prompt, context length, phrasing and position of the constraint, which suggests the loss is systematic. The same study found that a constraint-aware extractor running alongside the compactor raised retention above 90% without modifying the compactor or the model.

Two further properties make compaction hard to manage. Its output is unpredictable: the amount of output a summarizer produces, and the information it retains, can fluctuate substantially from run to run (Cim et al.). And the agent cannot easily tell what it has lost, because it treats the post-compaction context as complete.

What must never be summarized

The Lost in Compaction result suggests a general principle. Some context is compressible: its value degrades gracefully when summarized. Some is protected: summarizing it is a correctness failure. Ethen's internal architecture work lists the protected categories:

  • Obligations: what the task must still achieve, from the commitment graph.
  • Authority and permissions: what the agent may do, from its mandate.
  • Deadlines and time constraints.
  • Prerequisites for irreversible actions: conditions that must hold before money moves or data is deleted.
  • Approvals, bound to the exact actions they cover.
  • Active evidence: sources and quotations the current output depends on, verbatim.
  • Policy state: the policy version in force and any restrictions it imposes.

Compressible material includes completed work, superseded plans, old tool outputs, exploratory dead ends and conversational history whose decisions have already been recorded elsewhere.

Proposal: compile, do not just compress

The evidence-preserving context compiler assembles context through two channels (Figure 1).

Left: context sources (task state, commitment graph, mandate, conversation history, tool outputs, retrieved documents, completed work). Middle: a router splits them into a protected structured channel (obligations, authority and permissions, deadlines, irreversible-action prerequisites, approvals, active evidence, policy state) and a compressible channel. The compressible channel passes through a budget-aware compiler choosing full, summary, glimpse or exclude per item. Both channels merge into the assembled context. A rights filter applies before assembly.

Figure 1. Two channels: protected structure and compressible content. The compiler never summarizes the protected channel. Obligations, authority, deadlines, irreversible-action prerequisites, approvals and active evidence are carried as structured records. Only the compressible channel is subject to full / summary / glimpse / exclude decisions under the token budget. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen research proposal (evidence-preserving context compiler).

The protected channel carries the categories above as structured records, not prose. They are kept current by the systems that own them: the commitment graph, the mandate, the approval service. They are never passed through a summarizer. When the protected channel itself grows large, its records are pruned only by explicit state transitions: an obligation satisfied, an approval expired, a deadline passed. Text length is not a pruning criterion.

The compressible channel passes through a budget-aware compiler that chooses, for each item, one of four representations:

  • Full: the item verbatim.
  • Summary: a compressed version, with a link to the original.
  • Glimpse: a one-line reference that tells the agent the item exists and how to retrieve it.
  • Exclude: omitted entirely.

The choice is made under a token budget that reserves room for the model's output and tool results. Compaction is triggered when the projected context, including those reserves, would exceed a tested budget, rather than at a fixed percentage of the window. Ethen's internal design work proposes compacting at phase boundaries, such as after a sub-task completes, rather than mid-step. Compression methods for the compressible channel can draw on prompt-compression research, which has shown high compression ratios with limited loss on some tasks (Jiang et al., LLMLingua).

Relevance is learned; authorization is not. A learned model may decide which compressible items are relevant enough to include in full. It never decides what the agent is permitted to see. Permissions are enforced by a deterministic filter before compilation.

Rights propagate through compression

Compression creates derived artifacts: summaries, embeddings, cached contexts. Each inherits obligations from its sources. Ethen's internal architecture work states two rules:

  • Intersection, not union. A summary derived from several sources may be shown only to a caller who is permitted to see all of the sensitive sources that contributed to it, unless an explicit declassification rule applies. Combining a public document with a restricted one produces a restricted summary.
  • Permission-aware caching. Cache keys include a fingerprint of the permissions under which the cached context was assembled, so that a cached summary is never served to a caller with different permissions.

Revocation must propagate as well. When access to a source is revoked or a source is deleted, every summary, embedding and cached context derived from it must be invalidated, and re-validated at use time. A context strategy that keeps serving a summary of a document the user can no longer read has leaked that document. This is one reason rights must be infrastructure rather than policy text.

A residual risk deserves stating plainly. Even with correct permission filtering, an agent can sometimes infer restricted facts by combining individually permitted facts. Context engineering does not eliminate this inference leakage. It should be described honestly in security reviews rather than claimed away.

How this relates to memory systems

Agent memory systems address a related problem: what to remember across sessions. MemGPT pages information between a limited context and external storage, by analogy with operating-system virtual memory (Packer et al.). Mem0 extracts, consolidates and retrieves salient facts across conversations and compares this against full-context and retrieval baselines (Chhikara et al.). Zep uses a temporal knowledge graph that records when facts were valid, which helps with superseded information (Rasmussen et al.). Memory benchmarks are maturing. LongMemEval evaluates long-term interactive memory in chat assistants (Wu et al.). MemoryAgentBench identifies four competencies, namely accurate retrieval, test-time learning, long-range understanding and selective forgetting, and notes that existing benchmarks did not cover all four (Hu et al.).

The evidence-preserving compiler is complementary. Memory systems decide what to store and retrieve. The compiler decides how to assemble what is retrieved into a bounded context without losing obligations or violating permissions. A memory system can feed the compressible channel; it should not replace the protected one. A related distinction, between remembering documents and remembering how work gets done, is developed in Process Memory.

Research question and experiment

Research question (R04). Can task-conditioned context compilation reduce token cost while preserving obligations, permissions and evidence better than full context, fixed summarization, retrieval only or simple budget-aware selection?

Figure 2 summarizes our qualitative expectations for the five strategies. These are expectations to test, not results.

Matrix comparing full context, fixed summarization, retrieval only, budget-aware selection, and evidence-preserving compilation on: token cost, obligation retention, permission correctness after revocation, handling of stale or superseded facts, and predictability across runs. Full context is costly and degrades with length; summarization is cheap but drops constraints; the compiler is designed to retain obligations and permissions.

Figure 2. Context strategies compared on the properties that matter. Qualitative expectations for five strategies, to be tested. The proposal's claim is not that compilation always wins on tokens, but that it is the only strategy designed to hold obligation retention and permission correctness fixed while cost falls. Evidence label: QUALITATIVE MATRIX. Source: Ethen research proposal; expectations, not measurements.

Figure 3 shows the proposed design.

Experiment flow: task set seeded with probes (session constraints, deadlines, revoked permissions, superseded facts, deleted source documents) feeds five arms (full context, fixed summarization, retrieval only, budget-aware selection, proposed compiler). Each arm is scored on verified task success, obligation retention, revoked-context exposure, evidence accuracy, tokens and latency. Results compared at matched token budgets; probes measure retention while task success measures use.

Figure 3. Experiment design for research question R04. Five arms on the same tasks at matched budgets. Tasks are seeded with session constraints, deadlines, permission changes, superseded facts and deleted sources, and outcomes are scored on task success and on each preservation property separately. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

Tasks. Long tasks seeded with probes: session constraints stated early, deadlines, permissions revoked partway through, facts superseded by later information, and source documents deleted during the task. The VerifiedWork Context benchmark track defines a family of such tasks.

Arms. Full context (where it fits), fixed summarization, retrieval only, budget-aware selection without a protected channel, and the proposed compiler. All arms share the model, tools and task set, and are compared at matched token budgets.

Metrics. Verified task success; obligation retention, the share of seeded constraints honored at the end; revoked-context exposure, any appearance of content the caller could no longer access; evidence accuracy; tokens; and latency. Planted probes, which ask the agent about a constraint, measure retention, while task success measures use. Both are needed, because information can be present in context and still ignored.

Gate. Cost and latency gains with obligation retention and permission correctness at least as good as full context, under pre-specified margins. Kill conditions: hidden obligation loss on held-out task types, or gains that appear only on synthetic tasks and fail on independent real contexts. The protocol for estimating how much context different tasks actually need is How Should We Measure How Much Context an AI Agent Actually Needs?.

Failure modes

  • Protected-channel bloat. If too much is classified as protected, the protected channel itself exhausts the budget. Classification rules must be narrow, and pruning must follow state transitions promptly.
  • Misclassification. An obligation expressed casually in conversation may never be extracted into the protected channel. The constraint-aware extractor result suggests extraction can work well, but its error rate on real tasks must be measured.
  • Staleness. A protected record that is not updated when its source changes is worse than a summary, because it is trusted. Ownership of each record type must be explicit.
  • Overhead. Maintaining two channels and a compiler adds latency and complexity, which must be smaller than the savings.

Limitations

This proposal has not been tested. The 17% and 90% figures come from a single external evaluation with its own task distribution; Ethen has not measured retention on its own workloads. The protected categories are a design judgment. The compiler's value depends on obligations, permissions and evidence being available as structured records, which requires the surrounding infrastructure described in companion papers. Inference leakage across permitted facts is not addressed.

Conclusion

Context compression is unavoidable for long-running agents. What can be avoided is compressing the wrong things. By deciding what is protected before deciding how to compress, keeping obligations, authority, deadlines, approvals and evidence in structured form, and letting permissions flow through every derived artifact, an agent can spend fewer tokens without forgetting what it promised or seeing what it should not. Whether that holds on real work at acceptable overhead is the experiment we propose.

FAQ

Why not just use a bigger context window? Long contexts are costly, and models often use them poorly: benchmarks show large performance drops as length increases, even for models that accept very long inputs.

What is a session constraint? An instruction given once that should bind the rest of a session, such as "do not delete anything until I confirm". Current compactors frequently drop such constraints.

Does the compiler decide what the agent is allowed to see? No. Permissions are enforced by a deterministic filter before compilation. The compiler decides only how to represent permitted material.

References

  1. Wang, Z. et al. (2026). Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction. arXiv:2608.11242. https://arxiv.org/abs/2608.11242
  2. Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172. https://arxiv.org/abs/2307.03172
  3. Hsieh, C.-P. et al. (2024). RULER: What's the Real Context Size of Your Long-Context Language Models? arXiv:2404.06654. https://arxiv.org/abs/2404.06654
  4. Modarressi, A. et al. (2025). NoLiMa: Long-Context Evaluation Beyond Literal Matching. arXiv:2502.05167. https://arxiv.org/abs/2502.05167
  5. Cim, M. et al. (2026). Parallel Context Compaction for Long-Horizon LLM Agent Serving. arXiv:2605.23296. https://arxiv.org/abs/2605.23296
  6. Jiang, H. et al. (2023). LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. arXiv:2310.05736. https://arxiv.org/abs/2310.05736
  7. Packer, C. et al. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560. https://arxiv.org/abs/2310.08560
  8. Chhikara, P. et al. (2025). Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. arXiv:2504.19413. https://arxiv.org/abs/2504.19413
  9. Rasmussen, P. et al. (2025). Zep: A Temporal Knowledge Graph Architecture for Agent Memory. arXiv:2501.13956. https://arxiv.org/abs/2501.13956
  10. Wu, D. et al. (2024). LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. arXiv:2410.10813. https://arxiv.org/abs/2410.10813
  11. Hu, Y. et al. (2025). Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions. arXiv:2507.05257. https://arxiv.org/abs/2507.05257

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Product

    How to Explore Ethen Research Lab: Programs, Evidence Labels and Reading Paths

    The fastest way to read Ethen Research Lab well is to check two labels before reading anything else: the publication type (position paper, research note, proposal, technical report, methods paper, protocol, benchmark design, survey or system card) and the evidence status (measured result, synthesis, proposal, protocol, or external survey). Together they tell you what kind of claim the paper can make. Then filter the archive by research program to find papers on your topic, and use a reading path to follow a question from concept to benchmark to experiment. This guide explains each label, the programs, how related papers fit together, and where to start for your role.

  • Product

    What Is an AI Workspace? A Better Way to Think About It

    An AI workspace is a persistent place where people and AI systems work on something together. It holds six things a conversation does not: the context the work depends on, the artifacts being made, the state of the work (done, pending, blocked, unknown), the tools and permissions the AI may use, the people who own and review the work, and the history and evidence of what happened. A chat window can be one way to interact with a workspace, but it is not the workspace itself. The practical test is simple: if you closed the conversation, would the work still be there, organized, and checkable? If not, you were using a chat, not a workspace.

  • Product

    How We’re Rethinking AI Memory Across Ethen

    AI memory should not be one opaque pile of things an assistant decided to remember. Ethen's direction for memory rests on five ideas. Different kinds of memory get different rules: your preferences, facts about you, a project's context, your organization's knowledge, learned procedures and commitments you made each have their own scope, lifetime and controls. Memory is not the record of what happened: the authoritative record of a task's actions and outcomes is kept separately and never replaced by a summary. Memory carries its source and its validity, so it can be checked, superseded and explained. Permissions travel with memory, so a summary of a restricted document stays restricted and revoking access revokes what was derived from it. And you can see, correct, export and delete what Ethen remembers, and see when a memory influenced what Ethen did. This article explains each idea and what it means across Ethen's apps. It describes direction, not shipped architecture.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.