Skip to content

EthenEthenEthen

Research Note · 2026-10-03 · Trust & Accountable AI Work

Rights as Infrastructure: Building AI Datasets That Know How They May Be Used

Publication type
Research Note
Evidence status
Proposal / Hypothesis: A proposed direction, system, or hypothesis that remains untested.architecture proposal; requires legal review; no measured results
Research program
Trust & Accountable AI Work
Published
Authors
Ethen Research Lab
Reading time
12 min read

Most AI data carries its permissions in a contract somewhere else, if at all. We argue that every record and every derived artifact should carry its rights with it, so that a system cannot use data in a way its rights do not allow.

Cover image for "Rights as Infrastructure: Building AI Datasets That Know How They May Be Used". Decorative abstract motif; contains no data.

Abstract

Questions about AI training data rights are usually treated as legal review after the fact: a dataset is assembled, then someone asks whether it may be used. For AI systems that learn continuously from operational work, after-the-fact review does not scale and fails in predictable ways. Purposes are lost between capture and use. Consent is withdrawn after data has been copied into datasets, indexes and models. Derived artifacts such as summaries and embeddings escape the restrictions of their sources. This research note argues that rights should be treated as infrastructure. Every source record and every derived artifact carries a machine-readable rights record with independent dimensions: permitted purposes, consent state, sensitivity, residency, retention, contractual terms and lineage. Systems check those rights at the moment of use. We propose explicit semantic labels in place of overloaded class letters, and a separation of consent by purpose. We describe how revocation propagates through lineage to datasets and models, and why this design avoids unverifiable promises of machine unlearning. Nothing here is legal advice; every element requires review by counsel for any real deployment.

Why rights fail when they live outside the data

Consider an agent platform serving many organizations. A customer's support tickets are processed to resolve issues. Some of those tickets later become evaluation cases, some feed a routing model, and summaries of them are cached to speed up future work. Six months later the customer withdraws permission for any use beyond serving. Which datasets, evaluation sets, caches, embeddings and model versions are affected? In most systems nobody can answer quickly, because the purpose under which each record entered each artifact was never recorded with the artifact.

Several findings show how common this failure is in AI data generally. An audit of more than 1,800 text datasets found license omission rates above 70% and miscategorization rates above 50% on widely used hosting sites (Longpre et al., Data Provenance Initiative). A longitudinal audit of web domains used in training corpora found rapidly increasing restrictions on AI use over a single year, and inconsistencies between sites' terms of service and their crawler directives (Longpre et al., Consent in Crisis). Documentation practices such as datasheets for datasets (Gebru et al.) and model cards (Mitchell et al.) improve transparency, but they are documents, not enforcement.

The problem is sharper for operational data. Customer content in an AI product is almost always subject to contractual purpose restrictions. Detailed agent trajectories can identify the customer even after pseudonymization: one recent study showed that language models can link pseudonymous profiles at scale, reaching up to 68% recall at 90% precision in a closed-world setting where classical methods achieved near zero (Lermen et al.). "We removed the names" is not a rights basis for reuse.

Proposal: a rights record for every artifact

We propose that every source record, and every artifact derived from one, carry a rights record with independent dimensions (Figure 1).

Central box 'rights record for one source or derived artifact' surrounded by seven dimension boxes: permitted purposes (serving, tenant evaluation, tenant improvement, aggregate research, global training, public release); consent state and version; sensitivity (personal data, secrets, regulated); residency and processing location; retention and expiry; license or contract terms and rights holder; lineage (sources and descendants).

Figure 1. A rights record has independent dimensions. Collapsing these dimensions into a single class letter is the source of many errors. Each dimension can change independently: consent can be withdrawn while residency stays fixed; a license can expire while sensitivity does not change. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (semantic rights labels).

  • Permitted purposes. Which uses are allowed: serving the tenant, tenant-scoped evaluation, tenant-scoped improvement, cross-tenant aggregate research, global model training, public release. Purposes are explicit and enumerated, not inferred from a phrase like "to improve our services".
  • Consent state. Who granted the permission, under which version of which terms, when, and whether and how it can be withdrawn.
  • Sensitivity. Whether the content includes personal data, secrets, regulated categories or third-party confidential material.
  • Residency. Where the data may be stored and processed. For deployments where data may never leave the customer's environment, residency becomes a hard boundary for every downstream process; see Toward a Sovereign Improvement Protocol.
  • Retention and expiry. How long each purpose remains valid.
  • License or contract. The rights holder and the governing terms, for acquired data.
  • Lineage. Which sources an artifact was derived from, and which datasets, evaluation sets and models were derived from it.

Internal Ethen research recommends that the rights record be attached at capture, alongside the first telemetry for a task. The reason is simple: a purpose cannot be honestly backfilled onto data captured without one. Every day of untagged history is history that can never be used for anything beyond its original purpose.

Explicit labels instead of class letters

Earlier internal research used letter classes, A, B and C, for tenant-private, improvement-consented, and public, licensed or synthetic data. Different documents later used the same letters with different meanings, so the same letter meant tenant-private in one place and improvement-consented in another. A data system cannot tolerate that ambiguity. We propose explicit semantic labels instead (Figure 2), with residency and sensitivity carried as separate dimensions.

Matrix with six semantic labels as rows (tenant private, tenant improvement allowed, global improvement allowed, public licensed, expert owned, offline only) and six uses as columns (serve the tenant, tenant-scoped evaluation, tenant-scoped improvement, cross-tenant aggregate research, global model training, public release). Tenant private permits only serving and, where contracted, tenant evaluation; global improvement allowed permits training under its stated purposes; public release requires explicit permission in every case.

Figure 2. Semantic labels and the uses they permit. Explicit labels replace overloaded class letters. A label summarizes, but never replaces, the full rights record; residency and sensitivity are separate dimensions not shown here. The mapping is a proposal for discussion, not a legal determination. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal; not legal advice.

The labels are summaries for humans and dashboards. Enforcement always reads the full rights record. A label such as GLOBAL_IMPROVEMENT_ALLOWED does not permit every kind of global use. It permits the purposes listed in the specific grant, for its stated period.

A single checkbox for "data use" conflates decisions that people and organizations make differently. Internal research proposes separate, explicit grants for:

  1. operational processing, needed to deliver the service;
  2. tenant-local improvement, making the customer's own system better;
  3. de-identified or aggregate research, under stated privacy protections;
  4. global model training;
  5. public release, for example as part of a benchmark.

Grants should be specific, optional where possible, scoped in time and version, understandable, and withdrawable under defined terms. Many organizations will grant the first two and refuse the rest. Rights are one of the axes that determine whether data is an asset at all, as argued in What Makes AI Data Defensible?. A system designed so that most value comes from tenant-local improvement, as discussed in Private AI Improvement Without Raw Data Export, does not depend on obtaining broad grants.

Derived artifacts inherit restrictions

Rights must follow data through transformation. A summary of a restricted document is restricted. An embedding of personal data is personal data for most practical purposes. A synthetic example generated from private records is private-derived unless a documented review clears it. Three rules follow:

  • Intersection of sources. A derived artifact's permitted purposes are the intersection of its sources' permitted purposes, unless an explicit declassification rule applies.
  • Use-time checks. Rights are checked when an artifact is used, not only when it is created, so that later revocations take effect.
  • Lineage completeness. Every dataset, evaluation set and model version records the exact source rows and the purposes under which they were used. The build system that enforces this is the rights-aware dataset compiler.

These rules also govern the Work Receipt, which carries a rights record for each unit of agent work, and the Outcome Warehouse, which applies rights at query time.

Revocation and deletion

When consent is withdrawn or deletion is requested, the rights infrastructure makes propagation possible (Figure 3).

Sequence of boxes: revocation or deletion request; freeze new reuse; resolve descendants via lineage; remove or tombstone source copies; invalidate caches, indexes and summaries; rebuild affected dataset and evaluation versions; model response for affected trained artifacts (retrain, retire, or documented mitigation); verification; deletion receipt.

Figure 3. Revocation propagates through lineage. Withdrawing consent or deleting a source triggers a sequence that ends in a deletion receipt. Trained models are handled by retraining, retirement or a documented narrower remedy, never by an unverifiable claim that the model has forgotten. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (deletion propagation sequence).

The sequence follows internal Ethen architecture work. First, freeze new reuse. Then resolve descendants through lineage, remove or tombstone source copies, and invalidate caches, indexes and summaries. Next, rebuild affected dataset and evaluation versions and decide a response for each affected trained artifact. Finally, verify the result and issue a deletion receipt. Backups expire on their own schedule with tombstones recorded. Narrowly justified legal holds are documented.

Why we do not promise unlearning

The hardest case is a model already trained on data whose permission has been withdrawn. It is tempting to promise that the model will "unlearn" the data. The research literature does not support such a promise for large models in general.

Exact unlearning means producing the model that would have resulted without the data, typically by retraining. Sharded training schemes reduce the cost by limiting how much of the model any one record influences (Bourtoule et al.). Approximate unlearning adjusts parameters to approximate that result more cheaply. Several findings counsel caution. The definition underlying many approximate methods is problematic, because the same model can be obtained from different datasets, so closeness to a retrained model cannot by itself show forgetting (Thudi et al.). Unlearned language models can be "jogged" back into reproducing removed knowledge by relearning on small, loosely related data (Hu et al.). Benchmarks for unlearning in language models show how hard it is to establish equivalence with a model never trained on the data (Maini et al.). And security definitions that emulate perfect retraining can endanger the remaining data: an adversary controlling a few records may reconstruct much of a dataset through deletion requests (Cohen et al.).

We therefore propose a different commitment: retrain from lineage. Because every model version records exactly which data it was trained on, the system can identify affected models and retrain them from the remaining authorized lineage, retire them, or apply a documented narrower remedy. It does not claim that weights have forgotten anything. The simplest policy follows from the same logic: do not place uncertain customer content into globally trained weights in the first place.

Measuring whether rights infrastructure works

Rights infrastructure is only as good as its measured behavior, so we propose four operating measures.

Coverage. The share of records and derived artifacts that carry a complete rights record. Ethen's internal planning proposes requiring complete purpose and lineage records on every example that enters a curated dataset [PROPOSED TARGET]. Anything less means some data is being used on assumptions rather than recorded permissions.

Use-time denials. How often a use is refused because the rights record does not permit it. A rate of zero is suspicious. It may mean checks are not running, or that every request is pre-filtered so loosely that nothing reaches the check. Denials should be logged with their reasons and reviewed.

Revocation propagation time. The time from a revocation or deletion request to the point where no new use of the affected data, or of anything derived from it, can occur, and the separate time to complete rebuilds and model responses. These are different quantities and should be reported separately.

End-to-end deletion exercises. At least periodically, a synthetic record should be traced from capture through caches, indexes, datasets, evaluation sets and model lineage, then deleted, with every step verified and a deletion receipt produced. An exercise that cannot be completed reveals where lineage is missing before a real request does.

These measures turn rights from an assertion in a policy document into a property that can be audited. They also give buyers and regulators something more concrete than a promise.

Interoperability

The rights record should be expressible in open formats so that it can travel between systems. Provenance vocabularies such as W3C PROV-O can represent lineage. Dataset metadata formats such as Croissant can carry descriptive information. Neither format verifies copyright or consent, so the enforcement logic remains the system's responsibility. Content-provenance standards such as C2PA address a related problem for generated media outputs.

Rights infrastructure supports legal compliance; it does not establish it. Whether a given purpose requires consent, whether a contract permits a use, how data-protection law applies to embeddings or aggregates, and how obligations under regulations such as the GDPR's right to erasure or the EU AI Act apply to a specific system are questions for counsel. Recent litigation over training on curated copyrighted content shows that the legal landscape is moving. This note proposes engineering; it reaches no legal conclusions.

Limitations

The design has not been implemented end to end. Use-time checks add latency and complexity. Lineage completeness is hard to achieve retroactively for data already in use. The proposed labels and consent categories are illustrative and must be adapted to each organization's contracts and jurisdictions. Retraining from lineage is expensive for large models, which is one reason the design discourages placing uncertain data into them.

Conclusion

Rights that live in a contract and a spreadsheet will eventually be violated by a system that cannot see them. Rights recorded with the data, enforced at use, and propagated to every derived artifact make violations into errors a system can detect and refuse. They also make deletion something a system can actually carry out and evidence, instead of an unverifiable promise about what a model remembers.

FAQ

Why not just anonymize data before reuse? Because detailed records, including agent trajectories, can often be re-identified. Pseudonymization is not anonymization, and anonymity claims require analysis of the actual data.

Can a trained model be made to forget specific data? Not reliably for large models with current approximate methods. The proposed design retrains affected models from authorized lineage, retires them, or documents a narrower remedy.

Is this legal advice? No. Every element requires counsel review for a specific deployment and jurisdiction.

References

  1. Longpre, S. et al. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv:2310.16787. https://arxiv.org/abs/2310.16787
  2. Longpre, S. et al. (2024). Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv:2407.14933. https://arxiv.org/abs/2407.14933
  3. Gebru, T. et al. (2018). Datasheets for Datasets. arXiv:1803.09010. https://arxiv.org/abs/1803.09010
  4. Mitchell, M. et al. (2018). Model Cards for Model Reporting. arXiv:1810.03993. https://arxiv.org/abs/1810.03993
  5. Lermen, S. et al. (2026). Large-scale online deanonymization with LLMs. arXiv:2602.16800. https://arxiv.org/abs/2602.16800
  6. Bourtoule, L. et al. (2019). Machine Unlearning. arXiv:1912.03817. https://arxiv.org/abs/1912.03817
  7. Thudi, A. et al. (2021). On the Necessity of Auditable Algorithmic Definitions for Machine Unlearning. arXiv:2110.11891. https://arxiv.org/abs/2110.11891
  8. Hu, S. et al. (2024). Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning. arXiv:2406.13356. https://arxiv.org/abs/2406.13356
  9. Maini, P. et al. (2024). TOFU: A Task of Fictitious Unlearning for LLMs. arXiv:2401.06121. https://arxiv.org/abs/2401.06121
  10. Cohen, A. et al. (2026). Protecting the Undeleted in Machine Unlearning. arXiv:2602.16697. https://arxiv.org/abs/2602.16697
  11. W3C (2013). PROV-O: The PROV Ontology. https://www.w3.org/TR/prov-o/
  12. MLCommons. Croissant metadata format. https://mlcommons.org/working-groups/data/croissant/
  13. C2PA. Technical specification. https://c2pa.org/specifications/
  14. Regulation (EU) 2016/679 (General Data Protection Regulation), Art. 17. https://eur-lex.europa.eu/eli/reg/2016/679/oj
  15. Regulation (EU) 2024/1689 (Artificial Intelligence Act). https://eur-lex.europa.eu/eli/reg/2024/1689/oj

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Position Paper · Research Synthesis

    Verified Adaptive Intelligence: Learning From Work That Can Be Proven

    A research agenda for AI agents that learn only from experience that is verified, rights-cleared and shown to transfer across tasks, tools and model generations.

  • Research Note · Research Synthesis

    From AI Traces to Verified Experience

    Logs, traces, trajectories, outcomes and corrections are not the same asset. A research note on what turns agent telemetry into verified experience.

  • Position Paper · Research Synthesis

    What Makes AI Data Defensible?

    A position paper on the AI data moat: why volume is not defensibility, and eight axes, from rights to outcome density and transfer, that decide what compounds.

  • Product

    Ethen Studio: From Model Catalog to Creative Workspace

    Ethen Studio is evolving from a place to choose image, video, audio and voice models into a creative workspace: a place where a project, its history, its edits, its assets and its costs live together. A model catalog answers "which model can do this?". A creative workspace answers "how do I finish this project, keep it consistent, and know how each asset was made?". The catalog stays — Studio is the home of Ethen's full creative model catalog — but it becomes a tool inside the project rather than the whole product. This article explains what that shift means, what is already in place, what is direction, and what we are not claiming.

  • Product

    How We’re Rethinking AI Memory Across Ethen

    AI memory should not be one opaque pile of things an assistant decided to remember. Ethen's direction for memory rests on five ideas. Different kinds of memory get different rules: your preferences, facts about you, a project's context, your organization's knowledge, learned procedures and commitments you made each have their own scope, lifetime and controls. Memory is not the record of what happened: the authoritative record of a task's actions and outcomes is kept separately and never replaced by a summary. Memory carries its source and its validity, so it can be checked, superseded and explained. Permissions travel with memory, so a summary of a restricted document stays restricted and revoking access revokes what was derived from it. And you can see, correct, export and delete what Ethen remembers, and see when a memory influenced what Ethen did. This article explains each idea and what it means across Ethen's apps. It describes direction, not shipped architecture.

  • Product

    Why User-Controlled AI Memory Matters

    User-controlled AI memory matters because memory is what makes an AI assistant genuinely useful over time — and also what makes it personal, sensitive and capable of being wrong about you. An assistant that remembers your role, your projects and your preferences saves you from repeating yourself. The same memory can go stale, leak from one context into another, influence answers without your knowledge, or hold more than you meant to share. The answer is not to avoid memory but to put it under your control: you should be able to see what is remembered and where it came from, correct or delete it, pause remembering, keep personal and work memory separate, export it, and see when a memory influenced what the assistant did. Those controls are the direction for memory in Ethen.

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.