Skip to content

EthenEthenEthen

Research Note · 2026-10-03 · Data & Learning Systems

The Outcome Warehouse: Turning Completed AI Work Into Research Assets

Publication type
Research Note
Research program
Data & Learning Systems
Published
Authors
Ethen Research Lab
Reading time
12 min read

A task's outcome is not known when the task ends. It is revised by reopens, reverts, corrections and later evidence. A warehouse for AI work must treat outcomes as versioned facts with lineage and rights.

Cover image for "The Outcome Warehouse: Turning Completed AI Work Into Research Assets". Decorative abstract motif; contains no data.

Abstract

Organizations running AI agents need to answer questions that span many tasks. Which task families have the lowest cost per verified outcome? Which verifier versions disagree with expert reviewers? Which tasks regressed after the last model change? Which records are affected when a customer revokes permission? These questions require AI outcome data organized for analysis: outcomes that are verified, versioned over time, linked to the actors, models and verifiers that produced them, costed, and governed by the rights that apply to each record. This note proposes the Outcome Warehouse, an analytical store over Work Receipts and later evidence. It records outcomes as append-only, bitemporal facts. It separates execution, task and business success, and it enforces purpose restrictions at query time. We describe its data model, its handling of delayed and superseding outcomes, the questions it should answer, and its failure modes, including selection bias and the temptation to bill on weakly attributed business outcomes. The Outcome Warehouse is an Ethen architecture proposal that has not been built. It reports no measured results.

Why a separate store for outcomes

Operational systems answer operational questions: is this task running, did this call succeed, what does this customer owe this month. Research and improvement need different questions, asked across many tasks and over long periods:

  • Which configurations produce verified outcomes most cheaply, per task family?
  • How often do outcomes judged successful later reverse, and in which families?
  • Which verifiers disagree with experts, and has that changed across verifier versions?
  • Which failures recur after the fixes meant to prevent them?
  • If a customer withdraws permission for a dataset, which evaluation sets and trained artifacts are affected?

None of these can be answered reliably from logs, because logs lack outcomes. They cannot be answered from billing systems, which lack verification provenance, or from evaluation spreadsheets, which lack cost and rights. They require a store designed around outcomes as the central fact.

Internal research at Ethen ranked such a store among its most promising long-term assets, on the reasoning that verified outcomes may ultimately be more valuable than the system that generated them. That reasoning is a hypothesis. This note describes what the store would need to be for the hypothesis to have a chance.

Data model

Figure 1 shows the proposed architecture.

Left: inputs are work receipts, delayed evidence (reopens, reverts, acceptances, business signals) and human corrections. Middle: an append-only store of outcome versions with dimensions for task family, actors and model versions, verifier versions, cost and rights. A rights and purpose filter sits between the store and the consumers. Right: four consumers: research and evaluation builds, decision policy (Faros), economics (cost per verified outcome), and model change assurance. Tenant content stores are shown separately, linked by reference only.

Figure 1. Outcome Warehouse architecture. Receipts and later evidence arrive as append-only outcome versions. A rights filter applies at query time, not only at ingestion. Four families of consumers read governed views; none reads raw tenant content, which stays in tenant stores by reference. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (Outcome Warehouse).

The central entity is the outcome version: one judgment about one task's outcome, at one level of success, made by one verifier at one time. Its fields include:

  • Task reference. The task and its receipt, its family and difficulty class.
  • Level. Execution, task or business success (see below).
  • Verdict. Success, partial, failure, unknown, rejected, rolled back or superseded.
  • Verifier. Identity, version, verification level (deterministic, programmatic, expert or model judge) and calibrated error rates at the time of judgment.
  • Times. Valid time, when the outcome held in the world, and record time, when the warehouse learned of it. Keeping both makes the store bitemporal: it can answer "what did we believe on the first of the month?" as well as "what is true now?". Temporal knowledge graphs for agent memory adopt a similar bitemporal model for facts that change over time (Rasmussen et al.).
  • Supersession. A pointer to the version this one replaces, if any. Versions are never deleted or overwritten; deletion obligations are met by tombstoning and by removing content from tenant stores.

Dimensions link each outcome to the actors, model and tool versions, routing decisions, cost breakdown and rights record carried in the receipt. Corrections, the human-preferred alternatives captured at review time, are stored as linked facts with their own provenance. They are among the most informative records in the warehouse.

Two principles constrain the design. No raw content. The warehouse holds references and hashes, never prompts, documents or outputs, which stay in tenant-controlled stores under their own retention rules. Rights at query time. Every query declares a purpose: evaluation, decision-policy research, economics or training. The rights filter returns only rows whose grants permit that purpose at the time of the query. A grant that was valid when a record arrived but has since been revoked excludes the record from new queries.

Three levels of success

A recurring error in AI analytics is collapsing different kinds of success into one label. The warehouse keeps three (Figure 2), following the distinction used in Ethen's internal architecture work:

Timeline from left to right. At task end: execution success (calls returned, effects executed). After a reopen or revert window: task success (acceptance criteria met and not reversed). Weeks later: business success (retention, revenue, time saved), with a note that attribution to a single task is weak. Each level is drawn as a separate row with its own arrival time.

Figure 2. Three levels of success arrive at different times. Execution success is known at once; task success after a verification or reopen window; business success, if ever, weeks later and with weak attribution. The warehouse stores each as its own versioned fact rather than collapsing them into one label. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen architecture proposal (three success levels).

Execution success. The calls returned and the intended effects were executed. This is known immediately and says little about whether the task was accomplished.

Task success. The task's acceptance criteria were met, as judged by a verifier, and were not reversed within a defined window: a merge that was not reverted, a ticket that was not reopened, a refund that reconciled. This is the level most improvement work targets.

Business success. The work produced value: a customer retained, revenue recovered, time saved. This arrives weeks later, if at all, and its attribution to a single task is weak. It depends on many factors outside the agent's control.

Keeping the levels separate prevents two opposite mistakes. One is counting execution success as task success, which inflates performance. The other is billing or rewarding on business success, which hands the agent credit or blame for things it did not cause. Ethen's internal economic analysis cautions explicitly against billing broad business outcomes that depend heavily on customer actions and external conditions. The economic unit we propose instead is developed in Cost Per Verified Outcome.

Delayed outcomes and censoring

Because outcomes arrive late, a warehouse at any moment contains tasks whose task-success window has not closed. These tasks are censored: their final outcome is not yet known. Treating them as successes overstates performance. Dropping them biases the data toward fast-resolving tasks. The right treatment is to report them as pending and use methods that account for censoring when estimating rates. Every rate reported from the warehouse should state how many tasks were pending at the time of the report, and how the estimate handled them.

Late outcomes also change history. A task counted as a success last week may be reverted today. The bitemporal design lets an analyst reproduce last week's report exactly, which matters for audit. It also lets them see how much yesterday's numbers have since moved. That movement, the rate at which provisional successes flip, is itself a useful quality signal for each task family and each verifier.

Questions the warehouse should answer

Figure 3 lists representative questions and the fields each requires.

Matrix of six questions (cost per verified outcome by task family; which verifier versions disagree with experts; which tasks regressed after a model change; which failure classes recur after fixes; which outcomes later flipped; which datasets and models depend on a revoked record) against five field groups: outcome versions, verifier identity, cost, model and tool versions, rights and lineage.

Figure 3. Questions the warehouse should answer, and the fields they need. If a question cannot be answered from governed warehouse fields, the warehouse is incomplete for that purpose. Answering does not require raw content for any of these questions. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal.

Three of these deserve comment.

Cost per verified outcome by family needs outcome versions, verifier identity and cost, joined at the task-success level with pending tasks reported separately. It is the economic core of the decision-policy research in Faros, and it answers the call to report agent accuracy jointly with cost (Kapoor et al.).

Regression after a model change needs model versions on each outcome and enough historical outcomes per family to compare before and after. The warehouse supplies the task sample for Model Change Assurance.

Dependence on a revoked record needs lineage: which evaluation sets, training datasets and trained artifacts were built from which outcome rows. The warehouse records, for every dataset build, the exact rows it used and the purpose it declared. A revocation can then be traced forward to every affected artifact. The build process that maintains this lineage is the rights-aware dataset compiler.

From warehouse to research asset

Outcome rows become research assets only through deliberate builds. An evaluation set is a query over outcomes at a stated verification level, filtered by rights, split by task family and time, and sealed. A training set is a different query, with contamination controls that exclude any row derived from an evaluation item; canary strings and watermarking can make leakage detectable after the fact (Sander et al.). A decision-policy dataset joins outcomes with logged routing decisions and their propensities. Each build produces a versioned manifest that lists its rows and the rights under which they were used.

The warehouse is also the natural home for process memory: patterns in how tasks of a given kind actually unfold across systems, approvals and exceptions. That idea is developed in Process Memory. Which properties make the resulting assets defensible is the subject of What Makes AI Data Defensible?. What turns a single record into verified experience is covered in From AI Traces to Verified Experience.

Storage choices

The architecture does not require exotic infrastructure. Ethen's internal planning favors starting with a transactional database and object storage, adding analytical, streaming, vector or graph systems only when measured workloads justify them. The essential properties, append-only outcome versions, bitemporal times, query-time rights filtering and build manifests, can be implemented on conventional tools. Interoperable lineage can be expressed with a provenance vocabulary such as W3C PROV-O. Dataset documentation can follow established practices such as datasheets (Gebru et al.) and metadata formats such as Croissant. Audits of public dataset documentation find frequent license omissions and errors (Longpre et al.), which is one reason rights should be recorded per row at capture rather than inferred later.

Failure modes

Selection bias. The warehouse contains outcomes only for tasks that someone verified. If verification is applied more often to easy tasks, or only to tasks routed a certain way, the warehouse is biased in ways that propagate into every estimate. Verification sampling should be deliberate and recorded, so that estimates can be reweighted.

Goodhart pressure. Once warehouse metrics drive decisions, they attract optimization, in one of the several ways a measure can separate from its target (Manheim & Garrabrant). A team rewarded on verified-success rate may route hard tasks away from verification. Metrics should be paired with coverage measures, such as the share of tasks verified per family, that make this visible.

Label latency. Families with long reopen windows produce mature outcomes slowly. Decisions made on immature data favor fast-resolving work. Reports should show the maturity distribution of the outcomes behind every rate.

Privacy through metadata. Even without content, outcome rows reveal actors, timing, task types and resource identifiers. Metadata can identify customers and individuals, so warehouse access must be tenant-scoped by default, with cross-tenant aggregates released only under explicit grants and appropriate privacy protections.

Rights drift. Purpose grants expire and are revoked. A warehouse that applies rights only at ingestion will gradually serve rows it no longer has permission to use. Query-time filtering is the defense.

Limitations

The Outcome Warehouse is a design proposal. Its claimed value depends on verified outcomes accumulating in sufficient volume and diversity, which in turn depends on real deployments and on verifiers good enough to trust. Business-success attribution remains largely unsolved. Bitemporal designs are harder to operate than simple tables, and the operational cost has not been estimated. Questions of data ownership and permitted use require legal review for any real deployment.

Conclusion

An organization learns from its AI work only if it can ask what happened across many tasks, how it knows, what it cost, what changed later, and what it is allowed to do with the answer. The Outcome Warehouse is a proposal for asking those questions with care: outcomes as versioned facts, three levels of success kept apart, rights enforced when data is used, and every derived artifact traceable to the rows it came from.

FAQ

Is the Outcome Warehouse a data lake of agent logs? No. It stores outcomes and their provenance, not raw content or logs. Content stays in tenant stores, referenced by hash.

Why keep multiple versions of an outcome? Because outcomes change after tasks end. Keeping versions makes it possible to reproduce past reports, measure how often outcomes flip, and rebuild datasets correctly.

Can the warehouse measure business impact? It can store business signals when they arrive, but attribution to individual tasks is weak. We advise against billing or rewarding on business outcomes alone.

References

  1. W3C (2013). PROV-O: The PROV Ontology. https://www.w3.org/TR/prov-o/
  2. Gebru, T. et al. (2018). Datasheets for Datasets. arXiv:1803.09010. https://arxiv.org/abs/1803.09010
  3. MLCommons. Croissant metadata format. https://mlcommons.org/working-groups/data/croissant/
  4. Longpre, S. et al. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv:2310.16787. https://arxiv.org/abs/2310.16787
  5. Kapoor, S. et al. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
  6. Manheim, D., Garrabrant, S. (2018). Categorizing Variants of Goodhart's Law. arXiv:1803.04585. https://arxiv.org/abs/1803.04585
  7. Rasmussen, P. et al. (2025). Zep: A Temporal Knowledge Graph Architecture for Agent Memory. arXiv:2501.13956. https://arxiv.org/abs/2501.13956
  8. Sander, T. et al. (2025). Detecting Benchmark Contamination Through Watermarking. arXiv:2502.17259. https://arxiv.org/abs/2502.17259

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Position Paper · Research Synthesis

    Verified Adaptive Intelligence: Learning From Work That Can Be Proven

    A research agenda for AI agents that learn only from experience that is verified, rights-cleared and shown to transfer across tasks, tools and model generations.

  • Research Note · Research Synthesis

    From AI Traces to Verified Experience

    Logs, traces, trajectories, outcomes and corrections are not the same asset. A research note on what turns agent telemetry into verified experience.

  • Position Paper · Research Synthesis

    What Makes AI Data Defensible?

    A position paper on the AI data moat: why volume is not defensibility, and eight axes, from rights to outcome density and transfer, that decide what compounds.

  • Product

    What “Done” Should Mean for an AI Agent

    For an AI agent, "done" should mean that every requirement of the task has been met and that something other than the agent's own report shows it. Precisely: a task is complete when each required obligation is supported by evidence at the level of checking it needs — a passing test, a reconciled record, a confirmed delivery, an approved review — or has been explicitly waived by the person who owns the task. Three refinements make the definition usable. Keep execution success (the steps ran), task success (the outcome was achieved) and business success (it produced value) apart. Treat success as provisional until it can no longer be reversed. And report partial and unknown outcomes as what they are, instead of rounding them up to done.

  • Product

    How We’re Rethinking AI Memory Across Ethen

    AI memory should not be one opaque pile of things an assistant decided to remember. Ethen's direction for memory rests on five ideas. Different kinds of memory get different rules: your preferences, facts about you, a project's context, your organization's knowledge, learned procedures and commitments you made each have their own scope, lifetime and controls. Memory is not the record of what happened: the authoritative record of a task's actions and outcomes is kept separately and never replaced by a summary. Memory carries its source and its validity, so it can be checked, superseded and explained. Permissions travel with memory, so a summary of a restricted document stays restricted and revoking access revokes what was derived from it. And you can see, correct, export and delete what Ethen remembers, and see when a memory influenced what Ethen did. This article explains each idea and what it means across Ethen's apps. It describes direction, not shipped architecture.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.