Research Proposal · 2026-10-03 · Enterprise / Sovereign AI
Toward a Sovereign Improvement Protocol for Enterprise AI
If AI systems are to improve from the work they do, and the organizations they serve will not let that work leave, then improvement has to be redesigned around the boundary rather than across it. This paper sets out what such a design would need and how much of it is unsolved.
Abstract
Enterprises increasingly run AI systems inside environments they control: their own cloud accounts, dedicated regions, on-premises clusters and, in some sectors, disconnected networks. Sovereign AI in this sense means that the customer, not the vendor, decides where data lives, which models may process it and what may leave. That arrangement protects customers, and it cuts the vendor off from the evidence that would let its systems improve. This capstone paper of the Ethen research library asks how AI systems could improve across private deployments when raw customer data cannot leave the customer boundary. We propose the outline of a Sovereign Improvement Protocol: a layered design in which evidence and verification stay tenant-local, a rights and consent gate governs anything that leaves, only bounded statistics or updates cross the boundary under explicit privacy and security analysis, every release is signed, logged and auditable by the customer, and cross-tenant research happens only under rules agreed in advance. We describe each layer, distinguish what existing techniques provide from what is open research, set out the threat model, including malicious participants, and propose research phases with exit and stop criteria. The protocol is not production-ready and should not be read as a product description. It is a research agenda for a problem we consider unsolved.
The question
The previous papers in this library argue that verified experience, meaning outcomes checked by reliable verifiers, linked to the work that produced them and carrying their rights, is the most promising raw material for improving agents (Verified Adaptive Intelligence). The organizations that generate that experience are increasingly unwilling or unable to share it. Contracts prohibit vendor training on customer data. Regulation treats pseudonymized records as personal data where re-identification is possible, and large-scale re-identification from unstructured text is now practical (Lermen et al.). Some deployments are disconnected entirely.
So the research question is:
How can AI systems improve across private deployments when raw customer data cannot leave the customer boundary, and how can each customer verify what did leave and what it was used for?
The second clause matters as much as the first. A protocol that customers must take on trust is a contract, not a protocol.
What sovereign deployment means here
Figure 1 shows the deployment range the protocol must cover.
Figure 1. The deployment range a sovereign protocol must cover. Control over location, keys, model providers and egress shifts from vendor to customer along the range. The protocol must work at every point, including deployments where nothing leaves without a deliberate manual transfer. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen research proposal; deployment ladder synthesized from Ethen internal research.
At one end, a vendor hosts many tenants on shared infrastructure, with logical isolation. Next come dedicated single-tenant deployments operated by the vendor, then deployments inside the customer's own cloud account, sometimes called bring-your-own-cloud, where the customer holds the keys and the network controls. At the far end are on-premises and disconnected deployments, where nothing leaves without a deliberate, often manual, transfer.
Confidential-computing hardware can protect data in use from the infrastructure operator, and is useful when replay or learning runs on infrastructure the customer does not fully control. It is a technical control with its own threat model. It does not settle jurisdictional questions about who may lawfully compel access to data, and it does not decide what may leave the boundary. The protocol therefore treats it as one optional layer, not as the basis of sovereignty.
Design principles
- Local benefit first. Most improvement should come from learning that never leaves the tenant: its own retrieval, skills, routing and context policies, evaluated on its own work.
- Evidence and verification stay local. Raw content, per-task outcomes and verifier inputs never leave by default.
- Nothing leaves without a rights decision. Every outbound item passes a gate that checks purpose, consent and grant.
- Only bounded outputs cross. Outbound items are statistics or updates of a type, size and privacy cost fixed in advance.
- Every release is accountable. The customer can see exactly what left, when, under which grant, and how it was used.
- Revocation is honored forward. Withdrawing consent stops future use immediately and identifies everything derived from past contributions.
- Shared learning must earn its risk. Each step from local to shared must show measured utility that justifies its additional exposure.
The layered architecture
Figure 2 shows the proposed layers.
Figure 2. Seven layers of a Sovereign Improvement Protocol. Evidence, verification and local improvement stay inside the boundary. Anything that leaves passes a rights gate, is restricted to pre-declared bounded types, and is analyzed for privacy and security. Cross-tenant computation happens only in agreed studies, and benefit flows back locally first. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen research proposal (Sovereign Improvement Protocol).
Layer 1: tenant-local evidence
Work receipts, verified outcomes, failures, corrections and recoveries are recorded inside the tenant's boundary. This is the foundation, and it is the layer closest to existing practice: the evidence model is the same one described throughout this library, deployed where the customer controls it.
Layer 2: tenant-local verification and evaluation
Verifiers run inside the boundary. The tenant's own historical tasks become its evaluation set through Tenant Replay, so that model upgrades, skill changes and configuration changes are tested on the tenant's own work before they reach it. Model Change Assurance reports are produced locally and delivered to the customer. Customers can hold their own verifier keys and audit verifier behavior directly. Nothing at this layer requires any data to leave.
Layer 3: rights and consent gate
Anything that might leave passes a gate that resolves the rights of every contributing record: the purposes permitted, the grant version, its expiry and any revocation. The gate fails closed on unknown rights. Consent for improvement is separated by scope: operational use, tenant-local improvement, cross-tenant statistics, cross-tenant learning and public release are distinct permissions, and a grant for one does not imply another. The rights model is developed in Rights as Infrastructure, and the compilation procedure in A Rights-Aware Dataset Compiler.
Layer 4: bounded permitted outputs
Only outputs of pre-declared types may cross: per-family evaluation results for a candidate change, coarse configuration statistics, and, in research settings only, clipped and noised model updates. Each type has a fixed schema, a maximum size, a minimum cohort and a privacy cost. Free-form text, embeddings of customer content and skills derived from private work are not permitted output types, for the reasons set out in Private AI Improvement Without Raw Data Export.
Layer 5: privacy and security analysis
Each output type is analyzed before it is permitted. Where repeated releases are made, the tenant keeps a privacy budget ledger, with differential-privacy accounting at the tenant level where formal protection is claimed (Dwork & Roth; Abadi et al.), following published guidance for stating the unit of privacy and the accounting method (NIST SP 800-226). Where secure aggregation is used, the vendor learns only sums over many tenants (Bonawitz et al.). The analysis is recorded and reviewed by the customer's own privacy and security staff, not only the vendor's.
Layer 6: cross-tenant research under explicit rules
Cross-tenant computation, meaning federated evaluation, shared statistics or federated learning (McMahan et al.), happens only within a study whose purpose, participants, output types, privacy budget, duration and publication rules are agreed in advance by every participating tenant. The federated-learning literature is candid that robustness, privacy and fairness in these settings remain open problems (Kairouz et al.). The protocol treats every cross-tenant study as research until shown otherwise.
Layer 7: local and global benefit
Benefit flows back in two forms. Locally, every tenant receives improvements validated on its own work. Globally, where permitted, findings such as "this model upgrade tends to help on structured data entry" or "this recovery rule reduces duplicate effects" are shared with all participants and, if agreed, published. A tenant that contributes nothing still receives local benefit; a tenant that contributes receives evidence from others in return.
Accountability: signed releases and a transparency log
Figure 3 shows the accountability flow.
Figure 3. Accountability for every release. A release is signed inside the tenant, recorded in the tenant's own ledger, and appended to a public transparency log. The vendor records every use as provenance. Reconciling the three detects unauthorized releases, use beyond the grant and use after revocation. It does not by itself show that the vendor recorded everything. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen research proposal; mechanisms modeled on certificate transparency (RFC 9162), W3C PROV-O and in-toto.
Each outbound release is signed inside the tenant with a key the tenant controls, recorded in a local release ledger with the grant it was made under and its privacy cost, and submitted to an append-only, publicly verifiable log of the kind used for certificate transparency (RFC 9162). The vendor records every use of the release, including which aggregate, study or model it fed, as provenance records in a standard vocabulary (W3C PROV-O), and can attest to the build steps that consumed it (Torres-Arias et al.). A tenant can then compare its own ledger with the log and with the vendor's provenance records and detect any release it did not authorize, any use beyond the grant, or any consumption after revocation.
These mechanisms show that records are consistent and untampered. They do not show that the vendor recorded everything. Completeness depends on independent emission, reconciliation and audit, which is why the protocol includes third-party audit rights and why customers' own staff review the analysis.
Threat model
The protocol considers the adversaries described in Paper 39, with three additions specific to multi-tenant improvement.
Malicious participants. A tenant in a federated study can poison the shared result. In federated learning, a single participant can implant a backdoor by model replacement and evade anomaly detection (Bagdasaryan et al.). Defenses are participant admission, per-tenant contribution bounds, robust aggregation and, for evaluation studies, outlier review. They reduce rather than eliminate the risk, and they weaken with few participants.
Sybil participants. One party controlling several apparent tenants can amplify its influence or isolate a target tenant's contribution. Participant identity and admission must be verified outside the protocol.
Re-identification of tenants. Even when individual records are protected, an aggregate over a small group of tenants can reveal facts about one of them. Minimum cohorts at the tenant level, not only the record level, are required.
Vendor overreach. A vendor might use releases beyond their grant. The transparency log and provenance records make this detectable, and contracts make it actionable; neither makes it impossible.
Revocation
When a tenant withdraws a grant, future releases stop at once and no new study may use its past releases. Released aggregates cannot be recalled, which is why outputs are bounded and noised before release. Any model trained with the tenant's contributions is identified through lineage. Because exact removal of a contributor from a trained model is possible only for systems designed for it (Bourtoule et al.), the protocol's default is to keep revocable contributions out of shared weights entirely, and to retrain from lineage or retire a model if that rule was relaxed for a study. Where personal data is involved, erasure rights such as those in Article 17 of the GDPR apply in addition (Regulation (EU) 2016/679).
Regulatory and contractual uncertainty
Several questions that determine what the protocol may do are not settled. Whether a noised aggregate derived from personal data is itself personal data depends on jurisdiction and on the re-identification risk in context. Whether a model update counts as a transfer of the data that produced it is contested. Contracts that prohibit "training on customer data" rarely say whether federated evaluation, aggregate statistics or tenant-local learning are covered. Requirements for record-keeping and risk management in AI regulation and standards, such as the EU AI Act (Regulation (EU) 2024/1689) and the NIST AI Risk Management Framework, bear on how releases must be logged. This paper is not legal advice; each layer requires counsel review in each jurisdiction where it is deployed.
Utility against exposure
Not every output type is worth its risk. Figure 4 orders them.
Figure 4. Output types ordered by utility and exposure. Research proceeds down this table only when the row above has shown measured value. Free-form text, embeddings of customer content and skills derived from private work are not permitted output types at all. Evidence label: QUALITATIVE MATRIX. Source: Ethen research proposal; qualitative judgments, not measurements.
Per-family evaluation results for a defined candidate change carry high utility for little exposure: a tenant learns whether an upgrade is likely to help, and the vendor learns whether it helps across tenants, from a handful of numbers. Configuration statistics carry moderate utility and moderate exposure. Model updates carry potentially high utility and the highest exposure, and their utility for agents across enterprises is unmeasured. The research phases below follow this ordering.
Research phases
- Phase 0: local only. Tenant-local evidence, verification, replay and local improvement, with no cross-boundary flow. Exit criterion: measured local benefit on at least one task family. Stop criterion: none needed; this phase is useful alone.
- Phase 1: federated evaluation with synthetic tenants. The full protocol is exercised end to end using synthetic organizations, as proposed in Ethen Synthetic Enterprise, so that signing, logging, accounting and revocation can be tested with no real data at risk. Exit: every release reconciles between ledger, log and provenance; revocation drills pass.
- Phase 2: federated evaluation with consenting tenants. A small number of tenants evaluate defined candidate changes and release bounded, privatized results under explicit grants. Exit: results useful enough to change at least one decision. Stop: any unauthorized release, any re-identification finding in red-team review, or results too noisy to be useful at feasible cohort sizes.
- Phase 3: cross-tenant statistics. Configuration statistics under tenant-level accounting. Same exit and stop structure.
- Phase 4: cross-tenant updates, research only. Sealed studies with secure aggregation, tenant-level differential privacy and robust aggregation, using only data whose rights are not revocable within the study period. Proceeds only if Phases 2 and 3 show that shared evidence improves outcomes beyond local learning.
Every phase publishes its results, including decisions to stop.
What would show the protocol is not worth pursuing
The agenda should be abandoned or narrowed if local learning captures nearly all achievable benefit, so that shared evidence adds little; if privacy protection at feasible cohort sizes destroys the utility of shared statistics; if customers' privacy and security teams cannot be given enough visibility to trust the releases; or if the regulatory treatment of bounded outputs makes them impractical. Each of these is a plausible outcome.
Limitations
The protocol is a proposal and has not been built or tested. Several of its layers rely on techniques whose behavior for agent workloads, with long, heterogeneous trajectories and small tenant cohorts, is unmeasured. Its accountability mechanisms make misuse detectable but not impossible. Legal analysis is outside the scope of this paper and must be done per jurisdiction. The benefit of cross-tenant learning for enterprise agents is itself an open empirical question, and the protocol may turn out to matter mainly for evaluation rather than learning. Ethen has a commercial interest in the outcome, and the design should be reviewed by independent privacy and security researchers.
Conclusion
Sovereign deployment is becoming the expected arrangement for consequential enterprise AI, and it changes how improvement can work. A Sovereign Improvement Protocol would keep evidence and verification local, release only bounded and accountable outputs under explicit rights, treat cross-tenant learning as research, and let every customer check what left and how it was used. Much of the first layers can be built with known techniques. The later layers are open problems. Saying so clearly, and testing them in phases with stop criteria, is the most credible way to find out whether AI systems can improve from private work without asking anyone to give it up.
FAQ
Is the Sovereign Improvement Protocol available today? No. It is a research proposal. Tenant-local evaluation and improvement are the nearest-term layers; cross-tenant learning is research.
Does any raw customer data leave the boundary? Not under the protocol. Only pre-declared, bounded outputs may leave, under explicit grants, and each release is signed and logged.
Can a customer see what was shared? Yes, by design. The customer's own ledger, a public transparency log and the vendor's provenance records can be reconciled to show what left, under which grant, and how it was used.
Related research
- Private AI Improvement Without Raw Data Export — primitives and threat models.
- Tenant Replay: Private Evaluation Inside Enterprise Boundaries — tenant replay.
- Rights as Infrastructure: Building AI Datasets That Know How They May Be Used — rights infrastructure.
- Model Change Assurance: Testing AI Upgrades Before They Reach Real Work — assurance inside the boundary.
- Verified Adaptive Intelligence: Learning From Work That Can Be Proven — the adaptive-intelligence thesis at fleet scale.
References
- Lermen, S. et al. (2026). Large-scale online deanonymization with LLMs. arXiv:2602.16800. https://arxiv.org/abs/2602.16800
- Dwork, C., Roth, A. (2014). The Algorithmic Foundations of Differential Privacy. Foundations and Trends in Theoretical Computer Science 9(3–4). https://doi.org/10.1561/0400000042
- Abadi, M. et al. (2016). Deep Learning with Differential Privacy. arXiv:1607.00133. https://arxiv.org/abs/1607.00133
- NIST (2025). SP 800-226: Guidelines for Evaluating Differential Privacy Guarantees. https://doi.org/10.6028/NIST.SP.800-226
- Bonawitz, K. et al. (2016). Practical Secure Aggregation for Federated Learning on User-Held Data. arXiv:1611.04482. https://arxiv.org/abs/1611.04482
- McMahan, H. B. et al. (2016). Communication-Efficient Learning of Deep Networks from Decentralized Data. arXiv:1602.05629. https://arxiv.org/abs/1602.05629
- Kairouz, P. et al. (2019). Advances and Open Problems in Federated Learning. arXiv:1912.04977. https://arxiv.org/abs/1912.04977
- Laurie, B., Messeri, E., Stradling, R. (2021). RFC 9162: Certificate Transparency Version 2.0. https://www.rfc-editor.org/rfc/rfc9162
- W3C (2013). PROV-O: The PROV Ontology. https://www.w3.org/TR/prov-o/
- Torres-Arias, S. et al. (2019). in-toto: Providing farm-to-table guarantees for bits and bytes. USENIX Security 2019. https://www.usenix.org/conference/usenixsecurity19/presentation/torres-arias
- Bagdasaryan, E. et al. (2018). How To Backdoor Federated Learning. arXiv:1807.00459. https://arxiv.org/abs/1807.00459
- Bourtoule, L. et al. (2019). Machine Unlearning. arXiv:1912.03817. https://arxiv.org/abs/1912.03817
- Regulation (EU) 2016/679 (General Data Protection Regulation), Art. 17. https://eur-lex.europa.eu/eli/reg/2016/679/oj
- Regulation (EU) 2024/1689 (Artificial Intelligence Act). OJ L, 12.7.2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
- NIST (2023). AI Risk Management Framework 1.0 (NIST AI 100-1). https://doi.org/10.6028/NIST.AI.100-1
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Cost Per Verified Outcome: A Better Economic Unit for Agentic AI
A research note defining cost per verified outcome (CPVO) and comparing it with cost per token, request, seat and task as a unit for agentic AI economics.
- Ethen Synthetic Enterprise: An Executable World for Enterprise-Agent Research
A research proposal for enterprise agent simulation: an executable synthetic company with CRM, support, documents, identity, approvals, finance and email.
- Process Memory: Learning How Organizations Actually Get Work Done
A research proposal for process memory for AI agents: mining completed, verified work into per-organization process models that guide plans and flag anomalies.
Explained on the Ethen Blog
- How We’re Preparing Ethen for Private and Enterprise Deployments
Private AI deployment is a spectrum, not a single switch. At one end is a shared service with strong per-project controls; then a dedicated environment for one organization; then deployment inside the organization's own cloud account; and at the far end an isolated or offline environment that never touches an external network. Each step adds isolation, and each also changes things beyond where data lives: which models are available, who operates and updates the system, whether the system can improve from use, and how long it takes to start. Today, Ethen offers project-level controls — provider allowlists, budgets, logging modes including zero retention, customer-supplied keys — plus local models through Ethen Desktop and versioned GPU deployment recipes. Broader deployment options are direction, which we will pursue where demand warrants and describe only when they exist. This article explains the spectrum, the trade-offs, and why any label like "sovereign" must name its actual guarantees.
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.