Research Proposal · 2026-10-03 · Data & Learning Systems
A Rights-Aware Dataset Compiler for AI Training and Evaluation
A compiler refuses to produce a program that violates its type system. We propose the same discipline for datasets: a build that cannot prove every row is permitted for its purpose produces no dataset at all.
Abstract
Datasets for AI training and evaluation are usually assembled: rows are queried, filtered, transformed and exported, and questions about rights, contamination and lineage are answered afterward, if at all. This proposal treats dataset construction as compilation. A rights-aware dataset compiler takes a declarative build specification, consisting of a query over source records, a declared purpose, split rules and transformations. It resolves the lineage and rights of every candidate row at build time and checks purpose, residency, expiry and sensitivity. It runs contamination checks against sealed evaluation items, assigns splits by task family and time, and emits a dataset version with a signed manifest. Any unresolved question fails the build. Because every dataset version records exactly which rows and grants it used, a revocation can be recompiled into an exact list of affected datasets and models. That is the foundation for retraining from lineage rather than promising unlearning. We describe the compiler's phases and its error classes, its relation to existing data-documentation practice, its research questions and its failure modes, including the risk that strictness slows research. This is an Ethen architecture proposal. It has not been built, and its use for any real data requires legal review.
The problem: assembled datasets cannot be audited
A typical dataset build is a script. It queries a store, filters rows, applies transformations and writes files. The script may be versioned, but the dataset it produced usually is not tied to a record of which source rows it contained, under which permissions, and after which checks. Three practical failures follow.
Rights are lost. If a row's purpose restrictions are not checked when the dataset is built, a training set can contain data permitted only for serving. Audits of public datasets show how common incomplete rights metadata is: license omission above 70% and miscategorization above 50% on widely used hosting sites (Longpre et al.). Operational data adds contractual purposes that never appear in public metadata at all.
Contamination is invisible. If evaluation items and training rows share lineage, a model can be trained on its own test. The resulting scores look like capability and are actually memorization. Contamination of public benchmarks is well documented. A carefully constructed replacement for a grade-school math benchmark revealed accuracy drops for several model families (Zhang et al.). Watermarking and canary strings have been proposed to make such contamination detectable (Sander et al.). In both cases the damage happens at build time.
Deletion cannot be traced. When consent is withdrawn, the question "which datasets and models used this record?" has no answer unless every build recorded its rows. Without that record, honest deletion is impossible, and the remaining options are to over-delete everything or to under-delete and hope.
The companion note Rights as Infrastructure argues that every record should carry its rights. This proposal describes the build system that enforces them.
Proposal: compile datasets
A build specification declares:
- the query: which source records are candidates, for example verified outcomes of a given task family from the Outcome Warehouse;
- the purpose: what the dataset will be used for, such as tenant-scoped evaluation, cross-tenant aggregate research or global training;
- split rules: how rows are divided into training, development and sealed evaluation, by task family, source and time;
- transformations: redaction, aggregation, formatting or synthesis steps, each with its declared effect on sensitivity;
- exclusions: sealed evaluation sets and other lineages that must not appear.
The compiler processes the specification through fixed phases (Figure 1).
Figure 1. Compiler phases from build spec to signed dataset. A dataset build is a declarative specification compiled through checks in a fixed order. Any unresolved lineage or rights question fails the build; nothing is silently dropped or silently included. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (rights-aware dataset compiler).
- Resolve lineage. Every candidate row must trace to a source record. Rows that cannot be traced are errors, not omissions.
- Resolve rights. For each row, the compiler reads the rights record as it stands at build time, not at capture time. A grant revoked yesterday is respected today.
- Check purpose, residency and expiry. The declared purpose must be among the row's permitted purposes; the build's processing location must satisfy the row's residency; the grant must not have expired.
- Apply sensitivity transforms. Required transformations, such as removing personal data, are applied, and their effects are recorded. A transformation does not upgrade rights by itself. Redaction is not proof of anonymity, since language models can re-identify pseudonymous text at scale (Lermen et al.), and private-derived synthetic data remains private unless a documented review clears it.
- Run contamination checks. Rows derived from sealed evaluation items are excluded from training builds by lineage. Exclusion lists, canary-string scans, temporal cutoffs and near-duplicate detection catch leakage that lineage misses.
- Assign splits. Splits follow task family, source and time, not random rows, so that evaluation measures generalization rather than memorization of near-identical items.
- Emit a signed manifest. The dataset version is written together with a manifest listing every row identifier, the grant under which each was used, the transformations applied, the checks passed, the compiler version and a content hash.
An illustrative build
[ILLUSTRATIVE EXAMPLE — not an Ethen build.] A researcher wants an evaluation set of refund-handling tasks to compare two recovery policies across customers. The specification queries verified outcomes in the refund family from the past six months, declares the purpose cross-tenant aggregate research, splits by customer and month, and requires removal of personal names and account numbers.
The first compile fails. Lineage resolution finds that a batch of rows imported from an older system has no source identifiers; they cannot be traced, so they cannot be used. Rights resolution finds that most customers granted only tenant-scoped evaluation. Their rows are not permitted for cross-tenant research, and the compiler reports exactly which customers and how many rows. One customer's grant was revoked last week, so its rows fail even though they were permitted when captured. The contamination phase finds that twelve rows were derived from tasks already sealed in an existing benchmark.
The researcher now has three honest choices: narrow the purpose to a tenant-scoped evaluation for the customers whose grants permit it; request broader grants through the appropriate channel; or proceed with the smaller set of rows that compile, accepting the reduced coverage and reporting it. What the researcher cannot do is produce the original dataset by quietly ignoring the failures. The second compile, with a narrowed scope, succeeds and emits a manifest that records every exclusion and its reason.
Fail closed
The defining property is that the compiler fails closed. If any row's lineage or rights cannot be resolved, the build fails with row-level diagnostics. It does not drop the row silently and continue, because silent drops hide systematic problems: a missing grant table, a broken lineage join, a misconfigured purpose. It also does not include the row on the assumption that it is probably fine.
Figure 2 lists the error classes we propose.
Figure 2. Compile errors and their default behavior. Like a type checker, the compiler classifies problems and refuses to emit an artifact for any error. Overrides are possible only through an audited exception that is itself recorded in the manifest. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal.
Overrides exist, because real work sometimes needs them. An override is an audited exception with a named approver and a reason, and the override is recorded in the manifest. A dataset built with overrides says so on its face.
Reproducibility and lineage
Compiled datasets are reproducible: the same specification compiled against the same source versions and rights state produces the same content hash. That makes experiments auditable. A result reported on "dataset version 3" can be traced to the exact rows and permissions it used, and rebuilt if needed. Documentation practices such as datasheets (Gebru et al.) and metadata formats such as Croissant describe datasets for human readers and tools. The manifest is the machine-checkable counterpart that those descriptions can reference. Lineage can be expressed with a provenance vocabulary such as W3C PROV-O.
Revocation as recompilation
The strongest consequence of compiled datasets appears at revocation (Figure 3).
Figure 3. Recompiling after a revocation. Because every dataset version has a signed manifest of its rows and grants, a revocation can be compiled into an exact list of affected dataset versions, evaluation sets and model versions, plus the rebuilt datasets that replace them. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (retrain-from-lineage).
When a grant is revoked, the compiler re-runs every affected specification against the current rights state. Each recompile produces a new dataset version and a diff against the old one. The lineage index then lists every evaluation set and model version that consumed the old version. Each affected model receives a recorded decision: retrain on the new version, retire, or apply a documented narrower remedy. This is retrain-from-lineage. It replaces promises that a model has "unlearned" data, which current methods cannot substantiate for large models (Bourtoule et al.; Thudi et al.; Hu et al.), with an auditable procedure. The cost of retraining is one reason the surrounding policy discourages placing uncertain data into globally trained models at all.
Evaluation builds and training builds are different programs
The same source records may feed both evaluation and training, so the compiler must keep the two lineages apart. An evaluation build marks its rows as sealed; every subsequent training build excludes them and anything derived from them. Temporal cutoffs add a second barrier: training builds may include only records created before an evaluation item's creation time. Retired evaluation items may enter training only after a stated delay and a recorded governance decision. These rules make the measurement in Ethen VerifiedWork and the gain curves in How to Test Whether Verified Experience Improves AI Agents trustworthy. Neither is meaningful if training and test can mix.
Inputs for private aggregation
Some purposes involve no raw data leaving a tenant at all, for example computing aggregate statistics inside a customer's environment. The compiler applies there too. It decides which rows may contribute to an aggregate, enforces contribution limits, and records the privacy accounting that applies. The broader design for improvement without raw-data export is discussed in Private AI Improvement Without Raw Data Export.
Research questions
- Friction. How often do compile failures block legitimate research, and how much of that friction comes from missing metadata that could have been captured earlier?
- Granularity. Should lineage be tracked per row, per shard or per source? Finer granularity makes deletion precise but costs storage and build time.
- Purpose expression. Can permitted purposes be expressed in a form precise enough for mechanical checking yet close enough to contract language for counsel to review?
- Transformations and rights. Which transformations, such as aggregation with contribution limits or differentially private release, can legitimately change what a derived artifact may be used for, and how should the compiler represent that change?
- Recompile cost. What do recompilation and model response cost at realistic revocation rates?
Failure modes
- Strictness slows research. If builds fail frequently, researchers will route around the compiler. Fast diagnostics, clear ownership of rights metadata and a fast override path with audit are essential.
- False confidence. A dataset that compiles is consistent with its recorded rights. It is not thereby lawful, if the recorded rights were wrong. The compiler enforces what the rights records say and cannot correct them.
- Metadata debt. Historical data captured without rights records cannot compile into purposes beyond those clearly established. That is correct behavior, but it can make useful history unusable.
- Transformation laundering. A pipeline could try to "clean" restricted data through a sequence of transformations. Rights must flow through transformations by intersection unless an explicit, reviewed declassification applies.
Limitations
The compiler has not been built. Its practicality depends on rights records being captured with data from the start, which is not true of most existing corpora. Checking contamination by near-duplicate detection is imperfect. Purpose checking is only as good as the purpose taxonomy and the grants it reads. All use for real data requires counsel to confirm that the encoded rules reflect actual legal obligations and contracts.
Conclusion
A dataset is a program's input with legal and scientific consequences. Compiling datasets, so that rights, lineage, contamination and splits are checked before any artifact exists, turns those consequences into build errors instead of discoveries made after a model ships. It also makes deletion something a system can execute and evidence: recompile, diff, and decide for each affected model.
FAQ
Why fail instead of dropping problematic rows? Silent drops hide systematic problems, such as a broken lineage join or a missing grant table. Failing with diagnostics surfaces them.
Does a dataset that compiles comply with the law? Not by itself. It is consistent with the recorded rights. Whether those rights reflect legal obligations is a question for counsel.
How does this help with deletion? Every dataset version records its rows and grants, so a revocation can be recompiled into new versions plus an exact list of affected evaluation sets and models.
Related research
- Rights as Infrastructure: Building AI Datasets That Know How They May Be Used — the rights model it compiles.
- The Outcome Warehouse: Turning Completed AI Work Into Research Assets — the warehouse it queries.
- How to Test Whether Verified Experience Improves AI Agents — training builds for the gain-curve protocol.
- Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action — eval builds and contamination.
- Private AI Improvement Without Raw Data Export — private aggregation inputs.
References
- Longpre, S. et al. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv:2310.16787. https://arxiv.org/abs/2310.16787
- Zhang, H. et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332. https://arxiv.org/abs/2405.00332
- Sander, T. et al. (2025). Detecting Benchmark Contamination Through Watermarking. arXiv:2502.17259. https://arxiv.org/abs/2502.17259
- Gebru, T. et al. (2018). Datasheets for Datasets. arXiv:1803.09010. https://arxiv.org/abs/1803.09010
- MLCommons. Croissant metadata format. https://mlcommons.org/working-groups/data/croissant/
- W3C (2013). PROV-O: The PROV Ontology. https://www.w3.org/TR/prov-o/
- Bourtoule, L. et al. (2019). Machine Unlearning. arXiv:1912.03817. https://arxiv.org/abs/1912.03817
- Thudi, A. et al. (2021). On the Necessity of Auditable Algorithmic Definitions for Machine Unlearning. arXiv:2110.11891. https://arxiv.org/abs/2110.11891
- Hu, S. et al. (2024). Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning. arXiv:2406.13356. https://arxiv.org/abs/2406.13356
- Lermen, S. et al. (2026). Large-scale online deanonymization with LLMs. arXiv:2602.16800. https://arxiv.org/abs/2602.16800
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Verified Adaptive Intelligence: Learning From Work That Can Be Proven
A research agenda for AI agents that learn only from experience that is verified, rights-cleared and shown to transfer across tasks, tools and model generations.
- From AI Traces to Verified Experience
Logs, traces, trajectories, outcomes and corrections are not the same asset. A research note on what turns agent telemetry into verified experience.
- What Makes AI Data Defensible?
A position paper on the AI data moat: why volume is not defensibility, and eight axes, from rights to outcome density and transfer, that decide what compounds.
Explained on the Ethen Blog
- Why Not Everything Ethen Researches Needs to Become a Product
The relationship between research and product development at Ethen comes down to one rule: every research question ends in one of three outcomes — build, publish only, or stop — and only one of those is a product. Research becomes a product when five things line up: evidence from results rather than proposals, a real need from people doing real work, a cost of ownership we can sustain as models and data change, clearance on safety, privacy and rights, and a natural place in an existing product. Much valuable research meets some of those and not others. It may produce a method others can reuse, a benchmark, a safeguard inside Ethen, a design principle, or a negative result that saves everyone time. That is why a research publication from Ethen is never a product announcement, and why "publish only" and "stop" are normal outcomes rather than failures. This article explains the rule and how to read Ethen research with it in mind.
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.