Skip to content

EthenEthenEthen

Research Proposal · 2026-10-03 · Enterprise / Sovereign AI

Tenant Replay: Private Evaluation Inside Enterprise Boundaries

Publication type
Research Proposal
Evidence status
Proposal / Hypothesis: A proposed direction, system, or hypothesis that remains untested.research proposal; not built; requires security and legal review
Research program
Enterprise / Sovereign AI
Published
Authors
Ethen Research Lab
Reading time
13 min read

The most realistic test set an organization has is its own past work. It is also the one it can least afford to export. We propose evaluating on it where it already lives.

Cover image for "Tenant Replay: Private Evaluation Inside Enterprise Boundaries". Decorative abstract motif; contains no data.

Abstract

Enterprises evaluating AI agents face a tension. Public benchmarks are shareable but unrepresentative of their work. Their own historical tasks are perfectly representative but contain customer data, proprietary processes and regulated information that cannot leave their environment. Synthetic environments sit between the two and remain approximations. This proposal describes Tenant Replay: a form of private AI evaluation in which completed tasks are snapshotted, together with their recorded tool responses, and replayed under alternative configurations entirely inside the customer's boundary, with side-effecting tools stubbed at the network layer and verifiers running locally. Only bounded aggregates may leave, and only under explicit grants. Tenant Replay supports regression suites built from the customer's own work, model change assurance before a switch, and counterfactual evaluation of execution configurations, without raw-data export. We describe the architecture, a default policy for what may leave the boundary, how replay fidelity is measured per task family, the threat model, and the legal and security questions that must be answered before deployment. Tenant Replay is an Ethen architecture proposal. It has not been built.

The tension

Three sources of evaluation data are available to an organization deploying agents, and each fails in a different way.

Public benchmarks are reproducible and shareable, and they say little about an organization's own workflows. Contamination erodes their validity over time (Zhang et al.).

Synthetic environments can be controlled, shared and made hostile on purpose, as proposed in Ethen Synthetic Enterprise. They are approximations, and agents can learn their quirks.

The organization's own historical tasks are, by definition, its real distribution, including its rare cases, its systems and its policies. They are also its most sensitive data. Exporting them to a vendor for evaluation raises contractual, regulatory and security problems. Pseudonymization does not resolve them: detailed text records can be re-identified at scale with language models (Lermen et al.), and agent trajectories reveal customers through error strings, resource names, timing and tool sequences.

Tenant Replay resolves the tension by moving evaluation to the data rather than the data to evaluation.

Relation to existing evaluation practice

Tenant Replay borrows from three established lines of work and differs from each.

State-based agent benchmarks. Environments such as τ-bench (Yao et al.), AppWorld (Trivedi et al.), WorkArena (Drouin et al.) and CRMArena (Huang et al.) evaluate agents against controlled application state and check the resulting state rather than the transcript. Tenant Replay adopts the same principle, that outcomes are judged by state and verifiers rather than by plausible text, but the tasks come from one organization's completed work rather than from a curated public set. The cost is that every snapshot is private, so results cannot be compared across organizations in the way public leaderboards are.

Live and refreshable benchmarks. Claw-Eval-Live separates a refreshable layer of workflow demand from reproducible, time-stamped release snapshots with fixed fixtures and graders (Li et al., 2026). Tenant Replay applies the same separation inside one boundary: the organization's ongoing work is the refreshable signal, and each snapshot is a fixed fixture.

Federated evaluation. The federated-learning literature treats evaluation on data that stays on client devices or in client silos as an open problem with its own privacy and heterogeneity difficulties (Kairouz et al.). Tenant Replay is a narrow, institution-level instance: one tenant, one boundary, and an explicit gate on what leaves. It inherits the warning from that literature that keeping raw data local does not by itself make derived statistics safe.

Architecture

Figure 1 shows the proposed architecture.

A dashed tenant boundary contains: work receipts and snapshots of completed tasks; a snapshot store with recorded tool responses; a replay runner executing tasks under alternative configurations; stubbed side-effecting tools enforced at the network layer; verifiers; and a results store. An export gate applies aggregation thresholds and privacy controls. Outside the boundary: aggregate metrics only, for model change assurance reports and decision-policy research. Model inference calls from replay go only to providers permitted by the tenant's policy.

Figure 1. Tenant Replay architecture. Snapshots, recorded tool responses, replay runners and verifiers all run inside the customer's boundary. Side-effecting tools are stubbed at the network layer, not by instruction. Only aggregate results, after an export policy check, cross the boundary. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (Tenant Replay).

Snapshots. When a task completes, its relevant initial state, task specification, recorded tool responses, user turns and verifier definitions are captured as a replayable snapshot and stored inside the customer's environment. The general replay mechanism, and its fidelity trade-offs when an alternative configuration takes actions the original did not, are discussed in Counterfactual Replay for AI Agents.

Replay runner. A runner, deployed in the customer's own cloud account or data center, executes tasks from snapshots under alternative configurations: a candidate model, a different tool set, a new skill version or a different recovery policy.

Stubbed effects. Side-effecting tools are replaced by stubs that return recorded or simulated responses. Stubbing is enforced at the network layer, by blocking egress to production endpoints from the replay environment, not by instructing the agent. A replay that accidentally refunds a customer would be a serious incident, so prevention must not depend on the model's cooperation.

Local verifiers. Outcomes are judged by verifiers running inside the boundary, with certified error rates, as described in How Should We Measure the Reliability of LLM Verifiers?.

Model calls. Replaying a task with a candidate model sends that task's content to the model. That is new processing of customer data and must follow the customer's policy on permitted model providers and regions. For the strictest deployments, candidate models must be hosted inside the boundary.

Export gate. Per-task results stay inside. Only aggregates may leave, and only through a gate that applies grants, minimum cohort sizes and privacy review.

What may leave the boundary

Figure 2 shows a default policy.

Table with six artifacts: snapshots and recorded responses; per-task replay outcomes; per-family aggregate metrics; regression suites built from tenant tasks; environment templates (structure without content); model change assurance report. Columns: default location and conditions for leaving. Snapshots and per-task outcomes never leave; aggregates may leave above minimum cohort sizes with privacy review; templates only with explicit grant after content removal and review.

Figure 2. What stays inside, what may leave. A default policy for tenant replay outputs. Anything in the right-hand column leaves only under an explicit grant and the stated protections; nothing else leaves at all. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal; requires legal and security review.

Snapshots, recorded responses and per-task outcomes never leave. Regression suites built from a customer's tasks belong to the customer. A Model Change Assurance report is delivered to the customer, who decides whether to share it.

Aggregate metrics, such as success rates by task family under different configurations, may leave only under an explicit grant, above a minimum cohort size, and after privacy review. Even aggregates can leak information. Membership-inference and reconstruction attacks show that statistics and models derived from data can reveal facts about individual records (Shokri et al.; Carlini et al.). Where aggregates are released regularly, formal privacy accounting such as differential privacy may be appropriate. NIST has published guidance on evaluating differential-privacy claims (NIST SP 800-226), building on the formal framework of Dwork & Roth. Choosing parameters is a per-deployment decision with a utility cost.

Environment templates, meaning the structure of a task family without its content (which systems were involved, which approval steps occurred, what kinds of failure arose), are potentially valuable for building better synthetic environments. They may leave only with an explicit grant, after content removal and a re-identification review, because structure can identify an organization too.

Uses

Regression suites. Every organization accumulates a private test suite of its own representative and rare tasks. When anything changes, whether a skill, a tool integration or a harness setting, the suite runs before the change reaches live work.

Model change assurance. Before switching models, the organization replays a stratified sample of its history under the candidate model and receives a report of improvements, regressions and costs per task family. The validation protocol for such reports is A Research Protocol for Model Change Assurance.

Configuration research. Replays under alternative configurations produce counterfactual outcomes that inform which models, tools and verification levels work best for this organization's tasks, without exporting the tasks.

Improvement without export. Tenant Replay produces exactly the kind of local metric that privacy-preserving improvement methods aggregate, as discussed in Private AI Improvement Without Raw Data Export and Toward a Sovereign Improvement Protocol.

Fidelity, measured per family

Replay results are only as trustworthy as the replay (Figure 3).

Flow: select task family; measure stub coverage (share of tool calls with faithful recorded or simulated responses); replay the original configuration and compare verified outcomes with the original run (replay agreement); measure run-to-run variation as the noise floor; families passing pre-registered thresholds are eligible for assurance and research; others are reported as not replayable.

Figure 3. Measuring replay fidelity inside the boundary. Replay results are only as trustworthy as the replay. Two measurements bound that trust per task family: how much of the tool surface could be stubbed faithfully, and how often replay of the original configuration reproduces the original verified outcome. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen research protocol (proposed).

For each task family, two measurements bound that trust. Stub coverage is the share of tool calls in replayed runs that receive faithful recorded or simulated responses. Replay agreement is how often replaying the original configuration reproduces the original verified outcome. It is compared against the noise floor given by replaying the same configuration twice, since agents are stochastic. Families that fall below pre-registered thresholds are reported as not replayable, and conclusions about them are withheld. Hosted model behavior can change behind a stable name (Chen et al.), so model versions are recorded with every replay and agreement is re-checked periodically.

Families dominated by read operations and deterministic systems, such as code repositories in containers (Jimenez et al.) or state-based workflows (Trivedi et al.), are likely to replay well. Families whose outcomes depend on live human responses or external events are likely to replay poorly, and the design does not hide that.

Snapshot lifecycle and revocation

Snapshots are not permanent assets. Each carries the rights class and permitted purposes of the work it was captured from, and those can change.

Capture. A completed task is eligible for snapshotting only if its rights record permits tenant-local evaluation. Tasks whose outcome is still provisional, for example because a reopen or revert window has not closed, are captured but excluded from suites until the outcome settles, so that a regression suite does not encode a success that was later reversed.

Ageing. Systems change. A snapshot of a CRM workflow recorded against last year's schema may stop exercising anything the current system does. Snapshots therefore carry the versions of the tools and schemas they recorded, and suites report the share of snapshots whose recorded versions no longer match production. Old snapshots are retired on a schedule rather than accumulated indefinitely, because a suite dominated by stale tasks measures the past.

Revocation and deletion. When a customer revokes a purpose or a data subject's record must be deleted, every snapshot derived from the affected work is tombstoned, removed from regression suites and excluded from future aggregates. Aggregates already exported before revocation cannot be recalled. This is one reason the export gate is conservative: anything that leaves is harder to take back than anything that stays.

Lineage. Every replay result records which snapshots produced it, so that when a snapshot is withdrawn the affected results can be identified and, where needed, recomputed without it.

Competing explanations for a replay result

When a candidate configuration looks better or worse under replay, the difference has several possible causes besides a genuine change in capability, and each needs its own check.

  • Stub leniency. A stub that returns a recorded success for a call the candidate made differently can credit the candidate with an outcome it would not have achieved live. Divergent calls that fall outside the recorded set are counted, and families with high divergence are reported separately.
  • Path anchoring. Recorded responses come from the original configuration's path. A candidate that would have solved the task by a different route may be penalized because that route has no recorded responses. This biases replay toward configurations that behave like the original.
  • Verifier drift. If verifiers are updated between the original run and the replay, outcome differences may reflect the verifier rather than the agent. Verifier versions are pinned for each comparison.
  • Family selection. Families that pass fidelity thresholds are not a random sample of the organization's work. Conclusions are scoped to replayable families and are not extrapolated to the rest.
  • Run-to-run noise. Single replays of stochastic agents can differ by chance. Comparisons use repeated trials and the noise floor described above.

Threat model

We consider four adversaries. A curious vendor wants customer data: the architecture keeps raw data and per-task results inside the boundary, and the export gate limits what leaves. A compromised replay component might try to exfiltrate data or trigger real effects: egress controls and network-layer stubbing contain it, and the replay environment holds no production credentials. A malicious insider at the customer is outside the scope of what the design can prevent, beyond ordinary access controls and logging. An analyst combining exported aggregates over time might infer sensitive facts: minimum cohort sizes, release limits and, where justified, privacy accounting address this, without eliminating it.

Confidential-computing hardware can add protection when replay runs on infrastructure the customer does not fully control. Its protections hold only under its own threat model, and it does not settle jurisdictional questions about who may compel access to data.

Open research questions

  1. What stub coverage and replay agreement are achievable for common enterprise task families, and how quickly do they degrade as systems change?
  2. How many verified historical tasks does an organization need before replay-based assurance has useful statistical power for its main families?
  3. Can environment templates be released with a measured, acceptable re-identification risk, or does structure identify organizations too reliably?
  4. How well do replay-based predictions match live outcomes after a change is deployed? This is the validation question addressed by A Research Protocol for Model Change Assurance.

Limitations

Tenant Replay has not been built, and its fidelity and cost are unknown. Many customers will not have enough verified historical tasks at first. Deploying and operating replay infrastructure inside each customer's environment is expensive and depends on cooperation from their security teams. Replaying with candidate models raises data-processing questions that contracts may not yet address. All of the boundary and export rules described here require legal and security review for each deployment.

Conclusion

An organization's own past work is the best evidence of how a change will affect it, and the evidence it is least able to share. Tenant Replay proposes to bring replay, verification and comparison to where that evidence lives, export only bounded aggregates, and measure honestly where replay can and cannot be trusted. It would give enterprises assurance on their own terms, and it would give AI research an evaluation source that does not require anyone's data to move.

FAQ

Does Tenant Replay send customer data to Ethen? No. Snapshots, per-task results and regression suites stay inside the customer's boundary. Only aggregates may leave, under explicit grants and protections.

Can replay cause real side effects? It must not. Side-effecting tools are stubbed and production egress is blocked at the network layer, so prevention does not depend on the agent.

Which tasks can be replayed reliably? That is measured per task family through stub coverage and replay agreement. Families that fail are reported as not replayable.

References

  1. Zhang, H. et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332. https://arxiv.org/abs/2405.00332
  2. Lermen, S. et al. (2026). Large-scale online deanonymization with LLMs. arXiv:2602.16800. https://arxiv.org/abs/2602.16800
  3. Shokri, R. et al. (2016). Membership Inference Attacks against Machine Learning Models. arXiv:1610.05820. https://arxiv.org/abs/1610.05820
  4. Carlini, N. et al. (2020). Extracting Training Data from Large Language Models. arXiv:2012.07805. https://arxiv.org/abs/2012.07805
  5. NIST (2025). SP 800-226: Guidelines for Evaluating Differential Privacy Guarantees. https://doi.org/10.6028/NIST.SP.800-226
  6. Dwork, C., Roth, A. (2014). The Algorithmic Foundations of Differential Privacy. Foundations and Trends in Theoretical Computer Science 9(3–4). https://doi.org/10.1561/0400000042
  7. Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
  8. Jimenez, C. E. et al. (2023). SWE-bench. arXiv:2310.06770. https://arxiv.org/abs/2310.06770
  9. Trivedi, H. et al. (2024). AppWorld. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
  10. Drouin, A. et al. (2024). WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? arXiv:2403.07718. https://arxiv.org/abs/2403.07718
  11. Huang, K.-H. et al. (2024). CRMArena. arXiv:2411.02305. https://arxiv.org/abs/2411.02305
  12. Li, C. et al. (2026). Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows. arXiv:2604.28139. https://arxiv.org/abs/2604.28139
  13. Kairouz, P. et al. (2019). Advances and Open Problems in Federated Learning. arXiv:1912.04977. https://arxiv.org/abs/1912.04977
  14. Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Product

    Building Ethen Across Desktop, Web and Local AI

    Ethen is being built across three kinds of surface because no single one is best at everything. The web is where Ethen is quickest to reach: conversation in Ethen Chat and cloud workspaces you can open from any browser. The Ethen desktop app is where Ethen can work with your local files, your own toolchain and local coding. Local AI models, reached through the desktop app, let some work run entirely on your own machine — useful for privacy, offline work and experimentation. What ties them together is not that every surface does everything, but one account, shared project state where it makes sense, and the same rules about memory and approvals everywhere. Each surface and device keeps only the access it needs. This article explains that design. It describes product direction and published engineering; it does not announce downloads, release dates or capabilities beyond those already described in Ethen's engineering posts.

  • Product

    How We’re Preparing Ethen for Private and Enterprise Deployments

    Private AI deployment is a spectrum, not a single switch. At one end is a shared service with strong per-project controls; then a dedicated environment for one organization; then deployment inside the organization's own cloud account; and at the far end an isolated or offline environment that never touches an external network. Each step adds isolation, and each also changes things beyond where data lives: which models are available, who operates and updates the system, whether the system can improve from use, and how long it takes to start. Today, Ethen offers project-level controls — provider allowlists, budgets, logging modes including zero retention, customer-supplied keys — plus local models through Ethen Desktop and versioned GPU deployment recipes. Broader deployment options are direction, which we will pursue where demand warrants and describe only when they exist. This article explains the spectrum, the trade-offs, and why any label like "sovereign" must name its actual guarantees.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.