Skip to content

EthenEthenEthen

Position Paper · 2026-10-03 · Data & Learning Systems

What Makes AI Data Defensible?

Publication type
Position Paper
Research program
Data & Learning Systems
Published
Authors
Ethen Research Lab
Reading time
13 min read

The question is not how much data an organization has. It is whether that data can lawfully be used, is labeled correctly, would be hard for others to reproduce, and measurably improves work it has not seen.

Cover image for "What Makes AI Data Defensible?". Decorative abstract motif; contains no data.

Abstract

"Data moat" is among the most repeated claims in AI strategy and among the least examined. The usual version holds that an organization accumulating proprietary data, through users, deployments or partnerships, builds an advantage competitors cannot match. This position paper argues that the AI data moat, in its naive form, conflates volume with value. We propose eight axes along which data can be assessed: volume, label quality, uniqueness, rights, outcome density, verification provenance, reproducibility, and training utility, including transfer. We argue that defensibility depends mostly on the axes that are hardest to fake and slowest to accumulate. We review evidence that complicates naive accumulation: model collapse under recursive synthetic training, diminishing returns from scaling post-training data, widespread errors in dataset licensing metadata, rapidly increasing restrictions on web data, and the growing availability of open agent trajectories. We score common data assets qualitatively against the axes and propose a five-question screen for deciding whether a dataset is an asset at all. The paper is an Ethen Research Lab synthesis; it reports no measured Ethen results.

The naive moat

The naive argument runs like this. Foundation models are available to everyone. What differentiates one AI company from another is data that others do not have. Usage generates data, so the company with the most usage accumulates the most data and builds the strongest moat. Every link in this chain deserves scrutiny.

More data does not automatically improve models. Training language models on content generated by earlier models can cause them to lose the tails of the original distribution, a phenomenon called model collapse (Shumailov et al.). The severity depends on conditions, in particular whether synthetic data replaces or accumulates alongside real data, and the original claims have been debated (Seddik et al.; Borji). Separately, a systematic study of scaling reinforcement learning from human feedback found that gains from additional data, model size and method changes diminish along several dimensions (Hou et al.). Neither result says data is useless. Both say that more is not the same as better.

Usage data is mostly unlabeled. Most of what deployments generate records activity, not outcomes. As discussed in From AI Traces to Verified Experience, a log of what an agent did is not a record of whether it worked.

Some proprietary data cannot legally be used. Customer data usually comes with purpose restrictions. Web data is increasingly restricted: one longitudinal audit of 14,000 web domains found that, within a single year, about 5% or more of tokens in a widely used crawl corpus, and 28% or more of its most actively maintained critical sources, became fully restricted from AI use under robots.txt [EXTERNAL PRIMARY-SOURCE RESULT] (Longpre et al., Consent in Crisis). Even openly distributed datasets carry uncertain rights. An audit of more than 1,800 text datasets found license omission rates above 70% and miscategorization rates above 50% on widely used hosting sites [EXTERNAL PRIMARY-SOURCE RESULT] (Longpre et al., Data Provenance Initiative).

Generic data is becoming abundant. Open environment suites now release agent trajectories and trained verifiers alongside their tasks (Pan et al.). When anyone can generate or download comparable trajectories, trajectories alone confer little advantage.

Eight axes

If volume is not the measure, what is? We propose eight axes (Figure 1), each defined so it can be measured.

Table of eight axes: volume, label quality, uniqueness, rights, outcome density, verification provenance, reproducibility, and training utility or transfer. Columns give a definition and a measurement for each, for example label quality measured by verifier false-accept and false-reject rates and inter-rater agreement; uniqueness measured by whether a competitor could reproduce the data within a quarter from public sources and a frontier model.

Figure 1. Eight axes of data defensibility. Each axis is defined so that it can be measured. Volume is reported but never celebrated; a dataset whose training utility has not been measured is a candidate, not an asset. Axes adapted from Ethen's internal data-quality framework. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen internal synthesis (data-quality axes); proposed measurements.

Volume. Rows, tokens, episodes. It is easy to count and easy to celebrate, and it should be reported but never used as the headline.

Label quality. Whether outcome labels are correct. It is measured by the false-accept and false-reject rates of the verifiers that produced them and by agreement among human raters. See Evaluating the Evaluators.

Uniqueness. Whether a well-resourced competitor could reproduce the data from public sources and a frontier model within a fixed time and budget. This is a practical test, not a philosophical one. Synthetic data generated by prompting a public model is, by construction, reproducible by anyone with the same model.

Rights. The purposes for which the data may be used, by whom, until when, and whether permission can be withdrawn. Rights are not a compliance afterthought. A dataset that cannot be used for training is not a training asset, however large or unique. The infrastructure for recording rights is described in Rights as Infrastructure.

Outcome density. The share of records with a verified outcome attached, at a stated verification level. A small dataset with dense, verified outcomes can be more useful than a large one with none.

Verification provenance. Whether each label names the verifier, its version and its calibrated error, so that consumers can weight or exclude labels by trust tier.

Reproducibility. Whether the dataset can be rebuilt exactly from a versioned manifest of sources and transformations. Without it, results that depend on the dataset cannot be audited, and deletion obligations cannot be met precisely. Structured documentation of how a dataset was built, as proposed in datasheets for datasets (Gebru et al.), is a precondition.

Training utility and transfer. Whether adding the data improves held-out work on a pinned model, and whether the improvement survives changes of task family, tool and model. This is the axis that ultimately matters. It is also the one most often assumed rather than measured. Ethen's internal data framework adopts a blunt rule: a dataset with unmeasured training utility is a candidate, not an asset. The protocol for measuring it is How to Test Whether Verified Experience Improves AI Agents.

The axes interact, and the interaction is closer to multiplication than addition. A dataset with excellent labels and no usable rights is worth little. A dataset with clear rights and unknown label quality is risky to train on. A unique, well-labeled dataset that produces no gain on held-out work is an interesting archive, not an advantage. Figure 2 expresses this as a screening sequence: each question can disqualify the data, and later strengths cannot compensate.

Sequential decision flow with five questions: May we use it for the intended purpose? Are its labels verified with known error? Could a competitor reproduce it within a quarter? Has adding it improved held-out work on a pinned model? Does the improvement survive a model change? Each 'no' exits to a box describing what the data is instead: a liability, raw material, a commodity, a candidate, a depreciating asset. All yes leads to defensible asset.

Figure 2. Five questions before calling data an asset. A screening sequence. A dataset that fails an early question cannot be rescued by excelling at a later one: unusable rights or unknown label quality make volume irrelevant. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis.

The weakest-link view explains why naive accumulation disappoints. Accumulation improves volume, and sometimes uniqueness. It does nothing for rights, label quality, outcome density or verified training utility unless deliberate work is done on each.

Scoring common assets

Figure 3 scores common data types against five of the axes. The marks are qualitative judgments about typical instances, not measurements, and particular datasets may differ.

Matrix with ten data types as rows (raw chat logs, generic agent traces, large synthetic corpora, public benchmark data, open trajectories, outcome-labeled work records, expert gold sets, human correction records, rare failure and recovery cases, cross-configuration counterfactual outcomes) and five axes as columns (volume, uniqueness, rights clarity, outcome density, durability under better models). Raw logs score high on volume and low elsewhere; expert gold sets and rare failure cases score low on volume and high on uniqueness and outcome density.

Figure 3. Common AI data assets scored against the axes. A qualitative assessment of common data types. Marks are judgments about typical instances, not measurements of any particular dataset. The pattern matters more than any single cell: the assets that score well on uniqueness and outcome density are small and expensive. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis; qualitative judgment.

A pattern emerges. The assets with the highest volume, such as raw chat logs, generic traces and large synthetic corpora, score poorly on uniqueness, outcome density or both. The assets that score well on uniqueness and outcome density, such as expert gold sets, human correction records, rare failures with verified recoveries and cross-configuration counterfactual outcomes, are small, expensive and slow to accumulate. That combination is what makes them defensible. They require work, expertise, real tasks and time, and they cannot be backfilled.

Three asset types deserve emphasis:

  • Outcome-labeled work records. Records of real tasks with verified outcomes, costs and rights, such as the Work Receipts stored in an Outcome Warehouse. Their value grows with the diversity and difficulty of the work, not merely its volume.
  • Rare failures and recoveries. Failures carry more information per byte than routine successes, especially when the critical step and a successful recovery are known. See Toward a Failure Genome of Software Agents.
  • Environments and verifiers. An executable environment with calibrated verifiers can generate labeled experience on demand. High-fidelity environments have been shown to support training that transfers beyond the training distribution: in one enterprise customer-support simulation, reinforcement learning improved held-out task pass rates and also improved out-of-distribution tool-use benchmarks [EXTERNAL PRIMARY-SOURCE RESULT] (Mehta et al.). The authors attribute transfer to task-centric world building, expert-authored rubrics and realistic workflows. These are properties of the environment, not of the volume of data it produced.

Measuring training utility in practice

Because training utility is the decisive axis, it deserves a concrete method. We propose measuring it as a gain curve. Fix a base model version, a training or adaptation recipe, and a sealed evaluation set of task families never seen in development. Then train or adapt on increasing amounts of the candidate data, for example doubling each time, and measure the change in verified success and in cost per verified outcome on the sealed set. A dataset has training utility if the curve rises reliably above a matched-cost control. The control is the same budget spent on a comparison dataset, such as unverified traces of the same tasks or public data of the same kind. Utility should be reported per task family, with confidence intervals, because gains concentrated in one family can hide flat curves elsewhere.

Two refinements matter. Transfer is measured by evaluating on families, tools or model versions held out entirely from the data, not just on held-out instances of familiar families. Persistence is measured by repeating the evaluation after a model change: if the gain disappears when the base model is upgraded, the data was compensating for a weakness that no longer exists. Neither refinement is common practice, and both are necessary to tell a durable asset from a temporary patch.

Acquiring defensible data deliberately

If the valuable assets are small and slow to accumulate, waiting for usage to produce them is a poor strategy. They can be built deliberately, before large deployments exist. Expert-authored scenarios supply rare failures and gold-set labels, under contracts that assign ownership and record the authors' inputs. Executable environments generate labeled experience on demand; see Ethen Synthetic Enterprise. Licensed domain corpora can be acquired with explicit AI-training and evaluation terms rather than scraped, with per-source records of permitted purposes, territory and expiry. Partnerships can supply structure without content, for example workflow shapes or schemas donated under explicit purpose grants.

Each route trades money for time. None removes the need to measure utility. A purchased dataset that does not move a gain curve is an expense, not an asset, however well documented its rights.

What survives better models

A data strategy should be stress-tested against the possibility that foundation models become much better and much cheaper. Some data loses value in that world. Prompt-response pairs that taught a weaker model a skill the stronger model already has become redundant. Data whose main use was distilling expensive models into cheap ones loses value as cheap models improve.

Other data gains value. Verified outcome records on real tasks remain the only way to know whether a better model actually does better on that work. Gold sets for verifiers remain necessary, and arguably become more necessary as models become better at producing plausible but wrong outputs. Rare failures persist at the frontier of task length and complexity, because better models attempt harder tasks. Rights histories cannot be acquired retroactively at any price. The argument is developed in Why Better Foundation Models May Make Evaluation More Valuable, Not Less.

Counterarguments

Scale still matters for pretraining. True, and this paper does not concern pretraining-scale corpora, where volume and diversity remain central. Our argument concerns the data an applied AI organization accumulates through deployment, which is unlikely to matter at pretraining scale.

Network effects. Some products improve as more users interact, for example by surfacing popular content. Those are product network effects. They are real, but they are not evidence that interaction data trains better models.

Distribution beats data. Often true. An organization with strong distribution and mediocre data may outcompete one with excellent data and weak distribution. Data defensibility is one source of durable advantage among several, and not always the most important.

Synthetic data can be unique. Synthetic data generated with private seeds, expert-authored specifications or proprietary environments can be hard to reproduce. Its uniqueness then comes from the seeds, specifications and environments, which is consistent with our argument.

Data rights questions, including copyright in training data, contractual purpose restrictions and privacy law, vary by jurisdiction and are actively litigated. A September 2026 appellate decision in the United States, reported (Ballard Spahr) as holding that training a competing legal-research tool on a competitor's copyrighted headnotes was not fair use, illustrates the risk of relying on scraped curated content. Reports describe the ruling's scope as limited to a non-generative product. This paper does not offer legal conclusions. Any use of the axes above for decisions about specific data requires counsel review.

Limitations

The eight axes and the qualitative scores are judgments informed by literature and internal synthesis, not measured findings. The uniqueness test is a heuristic whose outcome depends on the assumed competitor. Training utility is the decisive axis, and we have not measured it for any Ethen dataset. Some of the evidence cited concerns pretraining or general language modeling rather than agent experience. Whether its lessons transfer is itself uncertain.

Conclusion

Data becomes defensible when it can be used, its labels can be trusted, it is hard to reproduce, and it measurably improves work it has not seen, including after models change. Volume contributes to none of these by itself. The practical implication is to invest in rights, verification, outcome density and measured training utility, and to treat any dataset that has not passed those tests as a candidate rather than a moat.

FAQ

Is a large proprietary dataset a moat? Not by itself. It must also have usable rights, verified labels, uniqueness and measured training utility. Size contributes to none of these.

What is outcome density? The share of records that carry a verified outcome at a stated verification level. It measures how much of a dataset can support a learning claim.

Can synthetic data be defensible? Its defensibility comes from what was used to generate it, such as private specifications, expert seeds or proprietary environments, not from the generated volume.

References

  1. Shumailov, I. et al. (2023). The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv:2305.17493. https://arxiv.org/abs/2305.17493
  2. Seddik, M. E. A. et al. (2024). How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse. arXiv:2404.05090. https://arxiv.org/abs/2404.05090
  3. Borji, A. (2024). A Note on Shumailov et al. (2024). arXiv:2410.12954. https://arxiv.org/abs/2410.12954
  4. Hou, Z. et al. (2024). Does RLHF Scale? Exploring the Impacts From Data, Model, and Method. arXiv:2412.06000. https://arxiv.org/abs/2412.06000
  5. Longpre, S. et al. (2024). Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv:2407.14933. https://arxiv.org/abs/2407.14933
  6. Longpre, S. et al. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv:2310.16787. https://arxiv.org/abs/2310.16787
  7. Pan, J. et al. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv:2412.21139. https://arxiv.org/abs/2412.21139
  8. Mehta, S. et al. (2026). EnterpriseBench CoreCraft: Training Generalizable Agents on High-Fidelity RL Environments. arXiv:2602.16179. https://arxiv.org/abs/2602.16179
  9. Gebru, T. et al. (2018). Datasheets for Datasets. arXiv:1803.09010. https://arxiv.org/abs/1803.09010
  10. Ballard Spahr (2026). Third Circuit Addresses Fair Use in AI Training but Leaves Generative AI Questions Unresolved (report on Thomson Reuters v. Ross Intelligence). https://www.ballardspahr.com/insights/alerts-and-articles/2026/10/third-circuit-addresses-fair-use-in-ai-training-but-leaves-generative-ai-questions-unresolved

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.