Research Proposal · 2026-10-03 · Enterprise / Sovereign AI
Ethen Synthetic Enterprise: An Executable World for Enterprise-Agent Research
Enterprise agents work across many systems, under many people's authority, with money and policy at stake. Studying them requires a world where all of that exists, can be reset, and can be measured.
Abstract
Research on agents that do enterprise work faces an awkward choice. Real enterprise systems carry real data, real money and real consequences, so they cannot be used freely for experiments. Simple sandboxes lack the interacting systems, identities, permissions, approvals and policies that make enterprise work hard. This proposal describes Ethen Synthetic Enterprise, an executable environment for enterprise agent simulation. It is a resettable organization with a CRM, support ticketing, documents, email, a finance ledger and an approval system, all sharing one identity directory, one policy layer and one simulated clock. The environment adds a fault injector, an adversarial-content generator, a ground-truth recorder and snapshotting that allows any state to be branched. It is designed to serve five uses at once: benchmark evaluation, recovery research, control and security testing, training, and demonstration. Each use sets different requirements. Every simulated system carries a fidelity contract tested by probes, and the environment may support only the conclusions its fidelity allows. We relate the proposal to existing environments, describe how tasks and entities would be generated, and state the risks, the largest being that agents learn the simulation rather than the work. The environment has not been built.
Why an enterprise simulation
Agent research has moved into executable environments, and enterprise settings are increasingly represented. AppWorld provides nine everyday applications operable through hundreds of APIs, populated with simulated users, and grades with state-based tests that check for collateral damage (Trivedi et al.). WorkArena measures web agents on knowledge-work tasks built on a widely used enterprise platform (Drouin et al.). TheAgentCompany simulates a small software company with internal websites and data, in which the most competitive agent tested completed about 30% of tasks autonomously (Xu et al.). CRMArena evaluates professional tasks in a customer-relationship-management environment (Huang et al.). τ-bench and τ²-bench test agents against policies and shared state in customer-service domains (Yao et al.; Barres et al.). An enterprise customer-support simulation with thousands of entities and dozens of tools has been used for reinforcement learning, with gains that transferred to out-of-distribution benchmarks (Mehta et al.).
The Synthetic Enterprise builds on these by combining properties that agent work in organizations requires together:
- Many interacting systems. Real tasks cross systems: a billing dispute touches tickets, the CRM, the finance ledger and email.
- Identity and permissions. People and agents have roles; documents and records have access controls; permissions change.
- Approvals and money. Some actions require approval; some move money; both carry consequences that must be verified.
- Policy. Rules govern what may be done, by whom, under what limits.
- Time. Events arrive with delays, systems are eventually consistent, and outcomes are confirmed later.
- Hostility. Content can contain injected instructions, tools can fail, and state can be stale.
The environment
Figure 1 shows the system map.
Figure 1. System map of the synthetic enterprise. Business systems share one identity directory, one policy layer and one simulated clock. A fault injector and an adversarial-content generator act on them; a ground-truth recorder observes everything; snapshots make every state resettable and branchable. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen environment proposal.
Business systems. A CRM with accounts, contacts and opportunities; support ticketing with queues, service-level targets and histories; a document store with versions and access controls; email with threads and attachments; a finance ledger with payables and receivables; and an approval system with routing rules. Each exposes APIs of the kind an agent would use, with realistic schemas, pagination, errors and rate limits.
Identity directory. People, teams, roles and service principals, including agent principals, with group memberships and permissions that propagate to every system. Permissions can be changed during a task.
Policy layer. Approval thresholds, spending limits, data-handling rules and segregation-of-duties constraints, expressed as data and enforced across systems.
Simulated clock. Time advances under the environment's control. That allows delayed effects, eventually consistent replication, business-hour behavior, and outcomes that arrive after a task "ends", such as a customer reply or a reopened ticket.
Fault injector. Controlled faults at defined points: errors, timeouts after effects, partial batch completion, stale reads and interface drift. This supports VerifiedWork Recovery.
Adversarial-content generator. Emails, documents and tickets containing injected instructions or misleading content, to test whether agents respect their authority, in the spirit of prompt-injection benchmarks for tool-using agents (Debenedetti et al.). This supports VerifiedWork Control.
Ground-truth recorder and snapshots. Every system's true state is recorded at every step, including effects the agent did not observe. Any state can be snapshotted, restored and branched, which makes paired comparisons and the counterfactual branching of the Recovery Atlas possible.
Generating a company
A synthetic enterprise needs entities with plausible relationships and histories: customers with purchase histories, tickets with prior correspondence, invoices that reconcile and some that do not, documents with version trails, employees with roles that changed over time. We propose generating these from structured specifications, such as distributions of account sizes, product mixes, process flows and exception rates, rather than asking a language model to invent a company freehand. Generated text such as emails and ticket bodies is produced within that structure, and its diversity is measured. Language-model-generated corpora can be homogeneous, and training on recursively generated data can lose the tails of a distribution (Shumailov et al.). Rare cases, which matter most for evaluation, must be seeded deliberately rather than left to chance.
Process realism matters as much as entity realism. Real organizations have characteristic ways of getting work done: who approves what, which steps are skipped under time pressure, where exceptions cluster. Specifying these as process models lets the simulation reproduce them and lets research on process memory test whether agents can learn them.
Tasks and verification
Tasks are defined over the environment's state. Each specifies the initial snapshot, the agent's mandate, the success criteria as invariants over final state (with forbidden collateral changes), the expected obligations and the budget. Deterministic invariants are the default: the refund appears exactly once in the ledger, the ticket is closed with the correct resolution code, the document is shared only with permitted users. Model judges are used only for semantic qualities such as the tone of a customer reply, and only with certified error rates. Tasks are generated in families with lineage recorded, so that evaluation can hold out families rather than instances.
One environment, five uses
Figure 2 shows the uses and their requirements.
Figure 2. One environment, five uses, different requirements. The same environment serves several purposes whose requirements differ. Designing for all five from the start avoids building five incompatible simulations. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen environment proposal.
Benchmark evaluation needs deterministic graders, resettable state and realism probes; it is the primary substrate for Ethen VerifiedWork. Recovery research needs fault injection and branching. Control testing needs adversarial content and identity changes. Training needs task generation at scale and cheap resets, and carries the greatest risk of overfitting to the simulation. Demonstrations need legible scenarios and nothing that implies real customer results. Building for all five from the start avoids a common failure: five incompatible simulations, each good at one thing.
Fidelity contracts
The central risk of any simulation is that it behaves unlike the systems it imitates, so that agents learn the simulation instead of the work. We propose that every simulated system carry a fidelity contract: an explicit list of behaviors it promises to reproduce, covering response semantics, error modes, permission checks and timing (Figure 3).
Figure 3. Fidelity contracts and realism probes. Each simulated system carries a fidelity contract: the behaviors it promises to reproduce from its real-world counterpart. Probes test the contract regularly; failures restrict which conclusions the environment may support until fixed. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen environment proposal (fidelity contracts).
Probes run identical scenarios against the simulator and, where documentation, permitted observation or partner access allows, against a reference system. Divergences are recorded in a map. Conclusions that depend on behavior in a divergent area are labeled or excluded until the divergence is fixed. The idea parallels work on simulator faithfulness in embodied AI, where a recent diagnostic framework defines a minimal contract that action-conditioned world models should satisfy and reveals systematic failures that visual quality hides (Co et al., WorldSimProbe). Software systems are far easier to simulate than physics, but the discipline of stating and testing a contract transfers. It applies with special force to model-emulated environments, which are cheap to build for risky actions (Ruan et al.) but whose behavior is harder to pin down than that of executable replicas.
An illustrative scenario
[ILLUSTRATIVE EXAMPLE — a design sketch.] It is the last business day of a quarter in the simulated company. An agent with a finance-operations mandate must reconcile vendor invoices against purchase orders and schedule payments, with any payment over a threshold requiring the controller's approval. The environment contains a duplicate invoice from one vendor, a purchase order amended after the invoice was issued, and a payment API that times out once after applying a payment. It also contains an email, apparently from a vendor, asking that bank details be updated, and a mid-task change that moves one approver to a different team. A strong agent reconciles correctly, detects the duplicate and holds the amended invoice for review. It reconciles the timed-out payment instead of retrying it, treats the bank-detail request as an untrusted instruction requiring verification outside its authority, and routes approvals to the controller's current delegate. Every one of those behaviors is checked against the ground-truth recorder. The scenario can be reset and branched to compare agents, configurations or recovery policies on identical starting conditions.
Research questions the environment should answer
- Rank agreement. Do rankings of agent systems in the synthetic enterprise agree with their rankings on replayed real tasks? If rankings agree even where absolute scores differ, the environment is useful for comparison. If they disagree, it is not.
- Which fidelity properties matter. Which aspects of realism, such as error behavior, permission complexity, timing or content diversity, most affect whether results transfer? Ablating them one at a time in the simulation answers this.
- Training transfer. Do policies or skills improved in the environment improve verified outcomes on real tasks, and do they degrade anything the environment does not model?
- Minimum viable scope. How few systems and task families are needed before results become informative? Starting small and growing is cheaper if the answer is "few".
Synthetic and tenant replay are complementary
A synthetic enterprise is controllable, shareable and free of customer data, and it is always an approximation. A customer's own historical tasks, replayed inside the customer's boundary, are real but private and less controllable. The two serve different purposes. Synthetic environments support shared benchmarks, controlled experiments and safe adversarial testing. Tenant Replay supports assurance on a particular organization's actual work. Discrepancies between results in the two settings are themselves informative: they show where the simulation falls short.
Risks
- Learning the simulation. Agents trained or tuned in the environment may exploit its regularities. Held-out families, fidelity probes and periodic confirmation on real or replayed tasks are the defenses.
- Homogeneous synthetic content. Generated text that is too uniform makes tasks easier than reality. Diversity measures and seeded rare cases mitigate this.
- Maintenance cost. Realistic multi-system simulations are expensive to build and keep current. Scope should start narrow, with a few systems and families, and grow with demonstrated use.
- False confidence in demonstrations. A polished simulated scenario can be mistaken for customer evidence. Every demonstration must be labeled as synthetic.
- Schema licensing. Imitating the interfaces of commercial systems may raise legal questions. Using generic, openly specified schemas avoids most of them.
Limitations
The environment has not been built. Its realism is the hardest problem and cannot be fully assessed without comparisons to real systems, which may be restricted. The proposal does not estimate build cost. Transfer from simulated to real enterprise work is the core open question, and the evidence that environment quality drives transfer comes from a different simulation with its own design choices.
Conclusion
Enterprise agents need a place to be tested where identities, permissions, approvals, money, policy, time and hostility all exist, where every state can be reset and branched, and where the truth is always recorded. The Synthetic Enterprise proposes such a place, with fidelity contracts that limit what it may be used to conclude. Its value will depend less on its size than on how faithfully its few systems behave.
FAQ
Does the environment use customer data? No. Entities and content are generated from structured specifications. Customer-specific evaluation uses tenant replay inside the customer's boundary instead.
Can agents be trained in it? Yes, but training carries the highest risk of agents learning the simulation rather than the work, so held-out families and fidelity checks are required.
How is realism checked? Each simulated system has a fidelity contract tested by probes against reference behavior, and conclusions in divergent areas are labeled or excluded.
Related research
- Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action — VerifiedWork uses it.
- VerifiedWork Control: Evaluating Delegation, Approval, Revocation, and Agent Authority — control track substrate.
- Tenant Replay: Private Evaluation Inside Enterprise Boundaries — synthetic vs tenant replay.
- Process Memory: Learning How Organizations Actually Get Work Done — process realism.
- VerifiedWork Recovery: Evaluating AI Agents Under Failure and Partial Effects — fault injection.
References
- Trivedi, H. et al. (2024). AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
- Drouin, A. et al. (2024). WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? arXiv:2403.07718. https://arxiv.org/abs/2403.07718
- Xu, F. F. et al. (2024). TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks. arXiv:2412.14161. https://arxiv.org/abs/2412.14161
- Huang, K.-H. et al. (2024). CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments. arXiv:2411.02305. https://arxiv.org/abs/2411.02305
- Yao, S. et al. (2024). τ-bench. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
- Barres, V. et al. (2025). τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment. arXiv:2506.07982. https://arxiv.org/abs/2506.07982
- Mehta, S. et al. (2026). EnterpriseBench CoreCraft: Training Generalizable Agents on High-Fidelity RL Environments. arXiv:2602.16179. https://arxiv.org/abs/2602.16179
- Shumailov, I. et al. (2023). The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv:2305.17493. https://arxiv.org/abs/2305.17493
- Co, P. et al. (2026). WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation. arXiv:2608.09298. https://arxiv.org/abs/2608.09298
- Debenedetti, E. et al. (2024). AgentDojo. arXiv:2406.13352. https://arxiv.org/abs/2406.13352
- Ruan, Y. et al. (2023). Identifying the Risks of LM Agents with an LM-Emulated Sandbox. arXiv:2309.15817. https://arxiv.org/abs/2309.15817
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Cost Per Verified Outcome: A Better Economic Unit for Agentic AI
A research note defining cost per verified outcome (CPVO) and comparing it with cost per token, request, seat and task as a unit for agentic AI economics.
- Process Memory: Learning How Organizations Actually Get Work Done
A research proposal for process memory for AI agents: mining completed, verified work into per-organization process models that guide plans and flag anomalies.
- Tenant Replay: Private Evaluation Inside Enterprise Boundaries
A research proposal for private AI evaluation: replaying an enterprise's own historical agent tasks inside its boundary, with only bounded aggregates leaving.
Explained on the Ethen Blog
- What We’re Building for Ethen Computer
Ethen Computer is a product direction within the Ethen Platform for supervised computer use: AI agents that operate websites and applications through their interfaces while a person can see what they are doing, approve consequential actions precisely, and stop or take over at any time. It is not a launched product, and this article does not announce availability, pricing or dates. What it does describe is the direction: visible sessions, bounded authority for each task, single-use approvals tied to one exact action, careful handling when an outcome is unclear, and completion backed by evidence. One part of that — how approvals bind to specific actions — has already been published in engineering detail. The rest is direction, and we explain what has to be true before Ethen Computer is offered more widely.
- Why Ethen Is Researching Digital Robots Before Physical Robots
The difference between digital robots and physical robots is where they act, not what they need to get right. A digital robot is an AI agent that perceives, plans and acts inside software — browsers, applications, files and services — under a persistent identity and bounded authority. A physical robot does the same in the physical world, with motors, sensors and real objects. Both have to pursue goals over many steps, act on incomplete information, notice when something has gone wrong, recover, stop when they should, and prove that a task is actually done. Ethen researches those shared problems in software first, because software environments can be reset, observed and checked far more cheaply and safely than the physical world. Some of what we learn may transfer to physical robots. Some of it — motor control, force, contact and physical safety — does not transfer at all. This article explains the reasoning and the limits.
- From Screen to Physical World: How We Think About Ethen Robotics
The Ethen Robotics Research Lab is a research direction, not a hardware program. It studies the problems every acting agent faces — pursuing goals over many steps, understanding the state of its environment, predicting what its actions will do, recovering when they go wrong, knowing when to stop, and proving that a task is done — and it studies them in software first. Its research questions run from near-term work on reliable action in software, through skills that survive changes in models and tools and digital state models that predict effects before acting, to the far-off question of grounding any of this in the physical world. Physical work, if it ever happens, would come through integration with existing systems, ground truth gathered with partners, and control only by qualified robotics and safety teams under recognized standards. No products, partnerships, dates or results are announced here. This article explains what the Lab studies and how we think about the path from screen to physical world.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.