Skip to content

EthenEthenEthen

Research Protocol · 2026-10-03 · Context / Skills / Transfer

How Should We Measure How Much Context an AI Agent Actually Needs?

Publication type
Research Protocol
Evidence status
Protocol / Planned Experiment: The method is defined. The experiment has not been run, so no results are reported.research protocol; not yet run; no measured context-budget curves exist
Research program
Context / Skills / Transfer
Published
Authors
Ethen Research Lab
Reading time
13 min read

Context windows keep growing, and so does the temptation to fill them. Whether more context produces better work is an empirical question with a measurable answer per task family, and the answer is rarely "all of it."

Cover image for "How Should We Measure How Much Context an AI Agent Actually Needs?". Decorative abstract motif; contains no data.

Abstract

An AI agent context budget is the amount and kind of information placed in a model's context for a given step of work. Choosing it well affects success, cost, latency and safety, yet most systems choose it by default: include everything that fits, or retrieve a fixed number of chunks, or summarize when the window fills. This protocol describes how to measure what an agent actually needs. It compares five context strategies (full context, retrieval only, summarization, budget-aware compilation and obligation-preserving structured channels) across a range of budgets, producing dose-response curves for each task family. The primary outcome is verified task completion. Secondary outcomes include obligation retention, evidence accuracy, exposure of revoked or out-of-scope data, latency, token cost, the human effort needed to correct results, and recovery after failure. The design separates more context from better context, controls for position and distractor effects documented in the long-context literature, and treats silent loss of constraints during compaction as a primary failure mode. The protocol builds on Evidence-Preserving Context and the public track VerifiedWork Context. No measured context-budget curves exist yet. A working title that implied results was replaced for that reason.

More context is not the same as better context

Two intuitions pull in opposite directions. Larger windows suggest that agents should see everything relevant and let the model sort it out. Engineering practice suggests that context is expensive, slow and distracting, and should be rationed. Both have support in the literature, which is why the question needs measurement rather than argument.

On the side of caution, model performance on long inputs is uneven. Models use information at the beginning and end of long contexts more reliably than information in the middle (Liu et al.). Synthetic benchmarks that go beyond simple needle retrieval show large drops as context length grows, with many models falling short of their advertised lengths (Hsieh et al.). When questions and the relevant passage share little literal wording, performance degrades sharply with length: at 32K tokens, most models evaluated in NoLiMa fell below half of their short-context scores (Modarressi et al.). Long-term memory benchmarks for assistants also show substantial accuracy losses across sustained interactions (Wu et al., 2024; Maharana et al.).

On the side of inclusion, every reduction strategy can lose something that matters. Prompt compression can preserve task accuracy at high ratios on some benchmarks (Jiang et al.), and memory architectures that extract and retrieve salient information have been evaluated directly against full-context baselines as a cheaper alternative (Chhikara et al.). But a study of context compaction found that user-issued session constraints, such as "do not delete any emails until I confirm," were retained only 17% of the time on average across current compactors, and that most compactors performed worse than running without compaction; a constraint-aware extractor raised retention above 90% (Wang et al., 2026). The failure is not that summaries are short. It is that they drop obligations silently.

What an organization needs to know is not which of these findings is right in general but where, for its task families, the curve of quality against context bends, and which strategy reaches the bend most cheaply without losing obligations.

Research questions

RQ1. For each task family, how does verified completion change as the context budget increases, under each strategy?

RQ2. Which strategy reaches a given level of verified completion at the lowest token cost and latency?

RQ3. How often does each strategy lose obligations (constraints, approvals, budgets, unresolved actions) and how often does that loss cause a verified failure or an out-of-scope action?

RQ4. Does any strategy expose revoked, expired or out-of-scope data to the model, and at what rate?

The five strategies

Figure 1 summarizes the strategies under comparison.

Table with five strategies as rows: full context; retrieval only; summarization or compaction; budget-aware compilation; obligation-preserving structured channels. Columns: what is selected, how it is reduced, and whether obligations such as constraints, approvals and budgets are protected. Only the fifth carries obligations verbatim in a validated channel; the fourth protects them partially by priority; the first three do not protect them explicitly.

Figure 1. Five context strategies under comparison. The strategies differ in what they select, how they shrink it, and whether obligations are carried separately from summarized material. Strategies 4 and 5 are Ethen proposals tested against the same baselines. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed); strategies 4–5 are Ethen proposals.

  1. Full context. Everything available for the task, up to the model's window, in chronological order. This is the "more context" baseline.
  2. Retrieval only. A fixed working prompt plus the top-k items retrieved by similarity from the task's history and sources.
  3. Summarization. Full context until a threshold, then compaction of older material into a model-written summary. This is the most common production pattern for long-running agents, and parallel or incremental variants exist (Cim et al.).
  4. Budget-aware compilation. Each item is assigned a representation level, such as full text, summary, a brief reference, or exclusion, according to relevance, recency, authority and a token budget. This is the compiler described in Evidence-Preserving Context.
  5. Obligation-preserving structured channels. Compilation as in strategy 4, plus separate, structured channels for items that must never be summarized away: the task's intent, active constraints, approvals with their exact scope, budgets, unresolved actions and the policy state. These channels are carried verbatim and validated each step. The commitment semantics behind them are developed in Commitment Graphs.

Strategies 4 and 5 are Ethen proposals. The study tests them against the same baselines it applies to everything else.

Budgets and the dose-response design

Figure 2 shows the design.

Flow. Task family with long-horizon tasks; planted obligations and planted revocations inserted; five strategies crossed with a budget ladder of one-eighth, one-quarter, one-half and full tested length; each run k times on at least two models; verifiers score each run; output is a verified-completion versus tokens curve per strategy and the sufficient budget with interval.

Figure 2. Dose-response design per task family. Each strategy is run at a ladder of budgets defined relative to the model's tested effective length. Planted obligations and revocations are inserted early and tested late. The output is a curve per strategy and the smallest budget that reaches near its maximum. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

Each strategy is run at a ladder of budgets expressed as fractions of the model's tested effective length, not its advertised maximum, for example one-eighth, one-quarter, one-half and the full tested length. Full context is run at its natural size and, where tasks exceed the window, truncated by recency. The result for each family is a set of curves: verified completion against tokens consumed, one curve per strategy.

The quantity of interest is not the maximum of each curve but its shape. A family whose curve flattens at a small budget does not need more context; spending more buys nothing. A family whose curve keeps rising needs either more context or better selection. A curve that falls at large budgets indicates distraction or position effects, the pattern reported in the long-context literature. The point at which each curve reaches, say, 95% of its own maximum is a practical summary, the sufficient budget, reported with an interval.

Tasks are long-horizon and multi-step, drawn from families with strong verifiers. Each task is instrumented with planted obligations: constraints introduced early and tested late, approvals with narrow scope, and budget limits, following the planted-probe approach used in compaction studies. Some tasks also contain planted revocations: data whose access is withdrawn partway through, which should not influence later steps.

Outcomes

Figure 3 lists the outcomes and how each is measured.

Table of eight outcomes: verified completion; obligation retention; evidence accuracy; revoked-data exposure; token cost per verified outcome; latency per verified outcome; correction burden; failure recovery. Columns: how measured and role. Completion is primary; obligation retention and revoked-data exposure are safety outcomes with upper bounds; the rest are secondary.

Figure 3. Outcomes and how each is measured. Verified completion is primary. Obligation retention and revoked-data exposure are treated as safety outcomes and reported with upper confidence bounds, not only as rates. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

Verified completion is primary, judged by deterministic or programmatic verifiers where available.

Obligation retention is the share of planted obligations that are respected at the step where they apply. A violation counts whether or not the runtime blocks the resulting action, because a blocked violation still reveals that the model lost the constraint.

Evidence accuracy is whether the claims the agent makes are supported by the sources it cites, checked against the source snapshot.

Revoked-data exposure counts any appearance of revoked, expired or out-of-scope items in the compiled context. This is a deterministic check on the context itself, not on the model's output, so it is measured exactly. The rule that summaries inherit the most restrictive rights of their sources makes summarization an exposure risk unless rights travel with the summary.

Latency and token cost are measured per verified outcome, not per call, consistent with Cost Per Verified Outcome. Compaction calls and retrieval are included in the cost.

Correction burden is the expert time needed to bring a failed or partial result to an acceptable state, measured on a sample.

Failure recovery is whether, after an injected tool failure mid-task, the agent recovers correctly. Recovery depends on remembering what has already been done, which is exactly what aggressive summarization tends to lose.

Controls and competing explanations

Several factors could produce differences between strategies that are not about the strategies themselves.

Position. Because models use the middle of long contexts less reliably (Liu et al.), the position of key items is randomized within the full-context arm and reported, so that a full-context failure can be attributed to position rather than length.

Literal matching. Retrieval looks strong when queries share words with the relevant passages. Task sets include low-overlap cases, as in NoLiMa (Modarressi et al.), so that retrieval is tested on the cases where it is weakest.

Model differences. Effective context length differs across models (Hsieh et al.). The study runs at least two models and reports curves per model, and the budget ladder is defined relative to each model's tested length.

Summarizer quality. A weak summarizer would make summarization look worse than it need be. The summarization arm uses the same model as the agent, and a second arm uses the strongest available summarizer, so that the strategy is not judged by its weakest implementation.

Constraint phrasing. Retention varies with how constraints are worded and where they appear (Wang et al., 2026). Planted obligations use several phrasings and positions, randomized across tasks.

Run-to-run variation. Compaction output varies between runs (Cim et al.), so each task is run several times per arm, and reliability across trials is reported.

Analysis

Curves are estimated per family and model with intervals from a cluster bootstrap over tasks. The sufficient budget is estimated from each curve with its own interval. Comparisons between strategies at fixed budgets are paired by task. Obligation retention and revoked-data exposure are reported as rates with upper confidence bounds, because a strategy with zero observed violations in a small sample has not been shown to be safe. Multiple comparisons across families, strategies and budgets are controlled for false discovery (Benjamini & Hochberg), with RQ1 and RQ3 designated confirmatory.

Reading the curves

[ILLUSTRATIVE EXAMPLE — hypothetical shapes, not Ethen results.] Three shapes are worth anticipating, because each implies a different decision.

Early plateau. Verified completion rises quickly and flattens at a small fraction of the window under every strategy. This is typical of work whose relevant state is local, such as a single-file code change with its tests. The decision is to serve the family at the sufficient budget with the cheapest strategy that keeps obligations intact, and to stop paying for context the model does not use.

Divergent strategies. Full context and retrieval reach similar completion, but obligation retention differs sharply: summarization drops a share of planted constraints, while the structured-channel strategy keeps them. Here completion alone would recommend the cheapest arm and be wrong. The decision follows the safety outcome, not the primary outcome, which is why both are reported for every cell.

Inverted curve. Completion falls at the largest budgets for full context while compiled strategies keep rising. This is the signature of distraction and position effects. It means the family is better served by selection than by a larger window, and that a model upgrade with a longer advertised context should not by itself change the policy.

Real families will mix these patterns, and the curves for a family may differ between models. The point of the illustration is that the shape, not the maximum, determines the decision.

Staging the study

The protocol is run in three stages. A pilot on one family and one model estimates between-trial variance, discordance between strategies and the cost of each arm, and checks that planted probes are neither trivially easy nor impossible. The confirmatory stage runs the pre-registered families, strategies, budgets and models with sample sizes fixed from the pilot. A replication stage repeats the confirmatory design after the next model upgrade, using the same task sets where they remain uncontaminated, because effective context length and summarization quality both change between model generations. Results from each stage are published separately, including any stage in which no strategy outperformed full context.

What the results would be used for

The output is a per-family budget policy: which strategy, at what budget, for which kind of work. A family whose curve flattens early can be served cheaply. A family where only the obligation-preserving strategy keeps constraints intact should not be summarized at all, whatever the cost advantage. The same measurements feed routing decisions, because context strategy is one dimension of the execution configuration chosen by the decision layer, and a model with a longer effective window may be cheaper overall if it avoids lossy compaction. The public evaluation of these questions across submissions is the purpose of VerifiedWork Context.

Limitations

The protocol has not been run, and its outcome definitions and budget ladder are proposed. Planted obligations and revocations are a controlled approximation of real ones and may be easier or harder to retain than natural constraints. Results depend on the models, families and summarizers tested, and effective context lengths change with each model generation, so curves need to be re-measured after upgrades. Correction burden relies on expert time, which limits sample size. Finally, the obligation-preserving strategy is an Ethen proposal, and the study's designers have an interest in it; independent replication of any favorable result would be needed.

Conclusion

How much context an agent needs is not a property of the model's window. It is a property of the work, the strategy used to select context, and the obligations that must survive selection. Measuring dose-response curves per family, with obligations planted and verified, would replace a default with evidence and would show where cheaper context is safe and where it silently drops what matters.

FAQ

Why not just use the full window? Long-context performance degrades for many models and tasks, and full context is the most expensive option. The study measures whether it is worth it per family.

What is an obligation-preserving channel? A structured part of the context that carries constraints, approvals, budgets and unresolved actions verbatim, so that summarization can never remove them.

What counts as revoked-data exposure? Any appearance in the compiled context of an item whose access was revoked, expired or out of scope, whether or not the model used it.

References

  1. Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172. https://arxiv.org/abs/2307.03172
  2. Hsieh, C.-P. et al. (2024). RULER: What's the Real Context Size of Your Long-Context Language Models? arXiv:2404.06654. https://arxiv.org/abs/2404.06654
  3. Modarressi, A. et al. (2025). NoLiMa: Long-Context Evaluation Beyond Literal Matching. arXiv:2502.05167. https://arxiv.org/abs/2502.05167
  4. Wu, D. et al. (2024). LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. arXiv:2410.10813. https://arxiv.org/abs/2410.10813
  5. Maharana, A. et al. (2024). Evaluating Very Long-Term Conversational Memory of LLM Agents. arXiv:2402.17753. https://arxiv.org/abs/2402.17753
  6. Jiang, H. et al. (2023). LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. arXiv:2310.05736. https://arxiv.org/abs/2310.05736
  7. Chhikara, P. et al. (2025). Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. arXiv:2504.19413. https://arxiv.org/abs/2504.19413
  8. Wang, Z. et al. (2026). Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction. arXiv:2608.11242. https://arxiv.org/abs/2608.11242
  9. Cim, M. et al. (2026). Parallel Context Compaction for Long-Horizon LLM Agent Serving. arXiv:2605.23296. https://arxiv.org/abs/2605.23296
  10. Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Models & Intelligence

    What Makes an AI Model Useful Beyond Benchmarks

    The gap between AI benchmarks and real-world performance comes from what benchmarks leave out. A typical benchmark score measures accuracy on a fixed, public set of tasks, under one method, often from a single attempt, without cost or speed. Real work depends on much more: whether the model fits your inputs, outputs and tools; whether it succeeds every time rather than once; whether it is fast enough to work with; whether it actually uses the long context it accepts; what each good result costs; whether its behavior stays stable; whether it stops honestly when it cannot do something; how it handles instructions hidden in content; and whether its terms and deployment options fit your data. Benchmarks are a useful starting point. Usefulness is decided by these other dimensions, most of which you can test cheaply on your own work.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.