Skip to content

EthenEthenEthen

Why Ethen Research Lab Publishes Its Work in Public

Ethen Research Lab publishes its work in public so that our claims can be checked, not just read. Every publication states what kind of evidence it contains — a measured result, a research synthesis, a proposal, a protocol or a benchmark design — and the first library of 40 papers says plainly that none of them reports a new measured Ethen result. Publishing that way does four things: it holds our claims to the evidence we actually have, lets others inspect our methods, commits us to how a hypothesis will be tested before any data arrive, and keeps research clearly separate from product claims. We also say what we keep private and why.

Ethen Research Lab publishes its work in public so that our claims can be checked, not just read. Every publication states what kind of evidence it contains — a measured result, a research synthesis, a proposal, a protocol or a benchmark design — and the first library of 40 papers says plainly that none of them reports a new measured Ethen result. Publishing that way does four things: it holds our claims to the evidence we actually have, lets others inspect our methods, commits us to how a hypothesis will be tested before any data arrive, and keeps research clearly separate from product claims. We also say what we keep private and why.

Key takeaways

  • Labels before conclusions. Every Ethen Research Lab publication declares its publication type and evidence status. The first 40 papers are syntheses, proposals, protocols, benchmark designs and surveys.
  • Tests before results. Thirteen of the forty papers are protocols or benchmark designs: they specify how a claim will be tested, including what would count against it, before any study is run.
  • Negative results count. Our protocols commit to publishing results that weaken our own hypotheses.
  • Some things stay private, on purpose. Customer data, private evaluation sets and security-sensitive details are not published, and papers say when something is withheld.
  • Research is not a product claim. A research proposal describes something we intend to study, not something Ethen products already do.

What does Ethen Research Lab publish?

Ethen Research Lab is Upcube's public research publication program. It publishes position papers, research notes, research proposals, technical reports, methods papers, benchmark designs, research protocols, surveys and system cards at the Ethen Research Lab archive. The archive can be filtered by research program and by publication type.

The first library, Research Publications V1, contains 40 papers organized into series: verified adaptive intelligence; trust and accountable AI work; evaluation and verification; model intelligence and Faros; data and learning systems; context, skills and transfer; and enterprise and sovereign AI. Alongside them, the archive holds the AgentTrustBench system card, which reports a measurement on a single pinned build of an Ethen system and states that it supports no population or live-traffic claim.

Table of five evidence classes with meanings and counts: A measured Ethen result 0, B research synthesis 7, C proposal 16, D protocol or benchmark design 13, E external survey 4.
Figure 1. Every paper declares its evidence class. The first library contains no new measured Ethen results, and says so on every page.

The evidence profile in Figure 1 is unusual for a company research program, and it is deliberate. We could have waited until we had measured results and published only the papers that looked good. We chose instead to publish the questions, the designs and the tests first, labeled for exactly what they are.

Why publish research at all?

A company that builds AI products could keep its research internal. We publish because four things become possible only in public.

Public claims have to match public evidence

A claim that is published with its evidence status attached is harder to inflate. When a paper says "research proposal; untested" at the top, nobody — inside or outside the company — can later cite it as proof that something works. Writing for the public record forces calibration: every "we propose", "we hypothesize" and "this remains unverified" is a commitment that can be checked against the text years later.

This matters more in AI than in most fields, because the distance between a persuasive demo and a reliable system is large. Agent research in particular has a history of headline numbers that turn out to depend on weak graders or narrow test sets. An audit of agentic benchmarks found task-setup and reward-design flaws that could misstate measured performance substantially, which is one reason our own benchmark designs require every grader's error rates to be published alongside any score.

Others can check the method

The machine learning community has spent years building norms that make research checkable: code submission, reproducibility checklists, and community efforts to re-run published results (Pineau et al., 2020). Documentation formats such as model cards and datasheets for datasets make the assumptions behind models and data visible to the people who rely on them (Mitchell et al., 2018; Gebru et al., 2018). Publishing our methods in full — definitions, metrics, baselines, failure modes and limitations — is how we take part in that system rather than asking to be exempted from it.

Publishing also exposes our mistakes earlier. A protocol with a weak baseline or an unfair comparison is far cheaper to fix when a reviewer points it out before the study runs than after the result has been announced.

Tests are fixed before results arrive

Thirteen of the forty papers are protocols or benchmark designs. A protocol describes an experiment that has not yet been run: the hypotheses, the comparison conditions, the metrics, the statistical plan, the stopping rules and the decision rule. Some include proposed thresholds — the bar a new method must clear before it replaces a simpler one.

Publishing those details before any data exist is the point. It prevents a common failure in applied research: running an experiment, looking at the results, and then choosing the comparison or metric that makes the result look best. When the test is public first, the result has to answer the question that was actually asked. Where a working title implied results we did not have, we replaced it with a protocol title, so that even the headline cannot overstate the evidence.

Five boxes in sequence: Question, Protocol published, Study run, Result published including negative results, Labels updated.
Figure 2. Thirteen of the forty papers are protocols or benchmark designs: tests published before any result exists.

A shared vocabulary for a young field

Agent research still lacks agreed terms for some of its most important problems. Several of our papers exist mainly to define a problem precisely enough to measure it. Commitment Graphs defines false completion — an agent reporting success while a required obligation is unmet — in a way that can be counted. Unknown Effects in Autonomous AI Systems argues that a timeout after an action is an unknown outcome, not a failure and not permission to retry. Cost Per Verified Outcome proposes measuring the cost of work that is independently verified rather than the cost of tokens. Definitions like these are only useful if they are public, so that others can adopt, criticize or improve them.

What publishing in public asks of us

Publishing in public changes how research is written, not only where it appears. Ethen Research Lab works under a short set of editorial rules that exist because the work is public.

Every number carries a label. A quantitative statement is marked as a measured Ethen result, an external primary-source result, an illustrative example, a proposed target, a hypothesis, or an unknown. An invented example used to explain arithmetic can never be mistaken for a forecast, and a threshold we intend to test against can never be mistaken for a result we achieved.

Calibrated verbs. Papers say "we propose", "we hypothesize" and "existing evidence suggests". They avoid "proves" where evidence only suggests, "guarantees" for security or privacy, and "solves" for open problems. Superlatives about Ethen are not used.

Internal ideas are not dressed up as literature. Ideas that come from Ethen's own work are attributed as an Ethen research hypothesis or architecture proposal. They are never cited as if they were peer-reviewed findings, and internal documents do not appear in reference lists.

No charts of numbers we do not have. Because the first library contains no measured Ethen data, none of its figures plots numbers as if they were data. Comparisons use qualitative marks, and legends say they are judgments.

Limitations are mandatory. Every paper has a limitations section, and the library as a whole publishes its major limitations on its index page.

These rules make the papers slower to write and less exciting to skim. They also make them safer to cite, which is the point of publishing them.

Why publish proposals before results?

The strongest objection to our approach is simple: why publish ideas that have not been tested? Wouldn't it be more rigorous to wait?

We think the opposite is true for two reasons. First, the alternative is not "no claims until results". It is usually informal claims — in talks, decks and product pages — with no evidence label at all. A labeled proposal is more honest than an unlabeled assertion. Second, publishing the design first creates accountability for the result. When we eventually run a protocol, readers can compare what we said we would test with what we actually tested.

Our protocols also commit to outcomes we would not enjoy. The verified-experience learning protocol, which tests the central hypothesis of our research program, states in advance which results would weaken that hypothesis and commits to publishing them. The VerifiedWork benchmark design commits to publishing results where Ethen-built systems perform worse than alternatives with the same prominence as results where they perform better.

What we keep private, and why

Publishing in public does not mean publishing everything. Some material stays private, and our policy is to say so rather than to pretend the published version is complete.

  • Customer and user data. Nothing in the research library contains customer content, and no customer or partner is named. Examples use generic roles.
  • Private evaluation sets. A test that is published can leak into training data and stop measuring anything. Benchmark designs therefore pair a public slice for reproducibility with sealed sets for decisive results, managed by reviewers independent of the system being tested.
  • Security-sensitive details. Papers describe defensive principles — for example, that approvals should bind to exact actions — but not deployment-specific values or weaknesses.
  • Some engineering know-how. Like most companies, we keep certain implementation details private. Papers describe mechanisms at the level needed to understand and test the research question.

When a paper depends on something it cannot publish, it says what is withheld and why. Readers can then judge how much the result depends on material they cannot see.

What public research means for Ethen products

Research publication and product announcement are different things, and we keep them apart. A research proposal describes something Ethen Research Lab intends to study. It does not mean an Ethen product already does it, will do it, or does it well. When research informs a product, the product article says so and links back to the canonical paper with its evidence status intact. We explain this separation in Why Ethen Keeps Research Separate From Product Claims.

The connection that does exist is a shared commitment: the products are built around making AI work checkable, and the research asks how to measure whether that checking works. The flagship paper, Verified Adaptive Intelligence, frames the central question as whether experience whose outcomes are verified — and whose reuse is permitted — can produce improvements that hold up on new work. It is a research synthesis and agenda. Answering it requires the protocols to be run.

Limitations and risks of this approach

Publishing this way has costs, and we want to name them.

Author interest. Ethen designs, and would benefit from, several of the proposals our protocols test. Independent replication is part of each protocol's plan, but readers should weigh our interest when reading our papers.

No measured results yet. A library of designs is easier to produce than a library of results. Until protocols are run, the value of this program is in its questions and methods, not its findings.

Partial coverage. Our own coverage audit records that not every internal source behind the papers was read in full, and that some external work was characterized at the level of its abstract. Those limits are stated in the program's public documentation.

Misreading risk. A proposal can be misread as a feature. Evidence labels reduce that risk; they do not eliminate it. We repeat the label wherever a paper is cited, including in this article.

How to read our papers

If you are new to the archive, start with the evidence status on each page, then read the abstract and the limitations section before the body. Our reader's guide, How to Explore Ethen Research Lab, explains the programs, publication types and reading paths in more detail, and What We Learned Publishing 40 Research Papers at Once describes how the library was produced and checked.

Frequently asked questions

Are Ethen Research Lab papers peer reviewed? They are published by Ethen Research Lab under its own editorial standards, not through a journal's peer-review process. Each paper states its evidence status and limitations so readers can judge it directly.

Does any Ethen Research Lab paper report measured results? None of the first 40 papers reports a new measured Ethen result. The archive's one measured result is the AgentTrustBench system card, which covers a single pinned build and makes no population or live-traffic claim.

Will you publish negative results? Yes. Our protocols state in advance which outcomes would weaken our hypotheses and commit to publishing them.

Is Ethen Research Lab the same as Ethen Research? No. Ethen Research Lab is the public research publication program. Ethen Research is a product for investigating questions. Ethen Labs is the company's research organization, whose approved work is published through Ethen Research Lab (Ethen Labs).

References

  1. Pineau, J. et al. (2020). Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program). arXiv:2003.12206. https://arxiv.org/abs/2003.12206
  2. Mitchell, M. et al. (2018). Model Cards for Model Reporting. arXiv:1810.03993. https://arxiv.org/abs/1810.03993
  3. Gebru, T. et al. (2018). Datasheets for Datasets. arXiv:1803.09010. https://arxiv.org/abs/1803.09010
  4. Zhu, Y. et al. (2025). Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv:2507.02825. https://arxiv.org/abs/2507.02825
  5. Ethen Research Lab (2026). Research Publications V1 (40 papers) and AgentTrustBench system card. https://upcube.ai/resources/research
  6. Ethen Research Lab (2026). Verified Adaptive Intelligence: Learning From Work That Can Be Proven. Position paper; research synthesis. https://upcube.ai/resources/research/verified-adaptive-intelligence
  7. Ethen Research Lab (2026). Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action. Benchmark design; not yet run. https://upcube.ai/resources/research/verifiedwork-benchmark