Skip to content

EthenEthenEthen

What We Learned Publishing 40 Research Papers at Once

Releasing Ethen Research Lab's first 40 papers together taught us that the hard part of publishing research at volume is not writing — it is keeping every claim matched to its evidence. The lessons that mattered most were procedural. Decide each paper's evidence status before drafting, and let the title follow it. Cite only sources whose claims have been checked against the source itself. Automate the structural checks, but know exactly what automation cannot catch. Treat redaction as its own review pass. Route papers that touch law, privacy or security to the right reviewers. And write down honestly what was not read in full. This retrospective describes that workflow so other teams can borrow it.

Releasing Ethen Research Lab's first 40 papers together taught us that the hard part of publishing research at volume is not writing — it is keeping every claim matched to its evidence. The lessons that mattered most were procedural. Decide each paper's evidence status before drafting, and let the title follow it. Cite only sources whose claims have been checked against the source itself. Automate the structural checks, but know exactly what automation cannot catch. Treat redaction as its own review pass. Route papers that touch law, privacy or security to the right reviewers. And write down honestly what was not read in full. This retrospective describes that workflow so other teams can borrow it.

Key takeaways

  • Evidence status comes first. A paper's class — synthesis, proposal, protocol, survey — determined its title, verbs and figures.
  • Titles were changed when evidence was missing. Working titles that implied results were replaced with protocol titles.
  • Citations were verified, not collected. Sources entered a reference registry only after their claims were checked; numbers were attributed only when the source stated them.
  • Automation enforced consistency, not correctness. Structural checks caught missing sections and broken links; reading caught weak arguments.
  • Redaction and review routing were separate passes. Sensitive categories were withheld deliberately, and papers needing legal, privacy or security review were flagged rather than published as if cleared.

What did we actually publish?

The first Ethen Research Lab library is 40 papers totaling roughly 88,000 words of prose, drawing on about 150 unique external sources. Each paper sits in one or more of seven series, from verified adaptive intelligence to enterprise and sovereign AI. Thirteen are protocols or benchmark designs, sixteen are technical or architecture proposals, seven are research syntheses and four are surveys of external literature. None reports a new measured Ethen result, and every paper says so. The papers are in the Ethen Research Lab archive. Why we chose to publish this way is explained in Why Ethen Research Lab Publishes Its Work in Public.

Releasing them together, rather than one at a time over months, had one large advantage: the papers could reference one another coherently. A concept paper could point to the benchmark design that would test it and the protocol that would run the test. It also had one large risk: a mistake in the workflow would repeat forty times. Most of what we learned is about managing that risk.

Lesson 1: Classify the evidence before writing a word

The single most useful decision was to fix each paper's evidence class at the planning stage, before drafting. A paper planned as a protocol is written as a protocol: hypotheses, design, analysis plan, stopping and decision rules — and no results section. A paper planned as a proposal is written as a problem, a mechanism, an evaluation plan and risks.

Doing this first changes the writing at every level. Verbs follow the class: "we propose" and "we hypothesize", not "we show". Figures follow the class: a proposal gets a conceptual or architecture diagram, never a chart that looks like data. Even section headings follow the class. When we tried the reverse order in early drafts — write first, label later — the label had to fight the text, and the text usually won.

Six boxes in two rows: Read the sources, Classify evidence first, Draft to the label, Verify every citation, Automated checks, Redaction and review routing.
Figure 1. The order matters most at the start: evidence status is decided before a word of the paper is written.

Lesson 2: Let the title follow the evidence

Titles are the most quoted part of any paper and the easiest place to overclaim. Several working titles implied findings we did not have — the natural phrasing of a question we hoped to answer. Where that happened, we replaced the title with a protocol title. "How reliable are LLM verifiers?" became "How Should We Measure the Reliability of LLM Verifiers?". A planned "failure genome" became "Toward a Failure Genome of Software Agents", because no failure dataset had been measured.

The change costs a little curiosity in the headline. It removes a much larger risk: a title circulating on its own, without the paper's careful evidence label, and being read as a result.

Lesson 3: Verify citations against the source, not against memory

Every external reference in the library comes from a registry of sources whose abstracts were checked at the time of writing. A number was attributed to a source only if the number appears in that source's abstract or is a well-known headline result of that work. Sources that could not be confirmed went on a do-not-cite list, and the list was binding.

This is stricter than common practice, and it caught real problems. Plausible-sounding statistics circulate widely in AI writing, often detached from where they came from. A registry forces two questions for every citation: does this source exist and say this? and is this the source that should be cited for it? The same discipline applies to our own internal material: ideas from Ethen's internal research are attributed as an Ethen research hypothesis or architecture proposal, never cited as if they were external literature.

The cost is that some papers characterize external work at the level of its abstract rather than its full text, and we record that limitation publicly instead of implying deeper verification than we did.

Lesson 4: Automate the checks you can state precisely

With forty papers, manual consistency is impossible. We encoded the house rules as automated checks and ran them on every paper and then on the library as a whole.

At the paper level, checks covered the word range, required sections (including limitations), a minimum number of meaningful figures and references, a calibrated-language scan for words such as "proves" and "guarantees", description length for search snippets, and a minimum number of contextual links to related papers. At the library level, checks confirmed that every internal link resolved, that no paper was orphaned, that references were numbered and cited in order, and that figures appeared in sequence.

Two lessons came out of this. First, writing a rule precisely enough to automate it improved the rule. "Avoid overclaiming" is advice; a list of banned constructions plus a required evidence label is a test. Second, automation is only as good as what it can see.

Table of five automated checks with what each catches and what it cannot catch.
Figure 2. Automation made the library consistent. It did not make it correct; that still required reading.

Lesson 5: Treat redaction as its own pass

A research library built from a company's internal work will contain material that should not be public. We handled that in two ways: exclusion while drafting, and a separate redaction pass afterward.

The categories withheld were decided in advance: repository and infrastructure details; commercial planning such as pricing, budgets and contract values; internal scoring of the company's own assets; intellectual-property strategy; internal model and vendor preferences; competitor-specific claims; and statistics without primary support. No paper names a customer or partner; examples use generic roles such as "a billing-operations agent".

An automated scan looked for known internal terms, identifiers and credential-like strings across every paper and figure caption. We learned quickly that a scan matches only what it knows. It cannot detect a sensitive fact phrased in ordinary language, and it cannot detect aggregation risk: papers that are each harmless but together reveal more than any one of them. Manual review covered the papers most likely to carry that risk, and the parts of the library flagged for security review are recorded as still needing a full human read. Figure source lines were also checked, because internal document names had crept into a few captions even though the paper text never used them.

Lesson 6: Readiness is not one status

"Done" for a research paper turned out to have several meanings. A paper could be ready for editorial review but still need legal review because it discusses data rights; ready for review but needing a privacy and security read because it describes authorization; or ready as a text while being results-gated, meaning no result may be claimed until its study is run. We tracked those as separate flags rather than one status. Ten papers were routed for legal or rights review and six for privacy or security review, and the routing is part of the library's public documentation.

The general point applies well beyond research: when a single "done" label hides several different checks, someone eventually assumes all of them passed.

Lesson 7: Write down what you did not read

The source material behind the library was large. Top-level syntheses and decision documents were read in full; many detailed research tracks were represented through their syntheses; and external work was in places checked at the abstract level. We published those limits rather than claiming exhaustive review. A coverage audit that says "partial" is less impressive than one that says "complete". It is also the only kind that stays true when someone checks it.

Lesson 8: One visual system, and no fake data

Every figure uses one restrained visual system, and every figure carries an evidence badge: conceptual diagram, qualitative matrix, proposed architecture, experiment design, taxonomy, or "illustrative — not measured Ethen data". Because no paper contains measured Ethen data, no figure plots numbers as if they were data. Comparisons use qualitative marks, and the legend says they are judgments. The most practical visual lesson was mundane: generated diagrams need their own layout checks, because text that overflows a box is easy to miss at review size.

Lesson 9: One question per paper, organized into families

Forty papers on overlapping subjects can easily end up competing with one another — for readers' attention and for search. We avoided that by giving each paper exactly one primary question and one intent, and by organizing related papers into families. On a given topic, a concept paper owns "what is this and why does it matter?", a benchmark design owns "how would we compare systems?", and a protocol owns "how would we test this specific claim?". Where two papers used nearly the same phrase — an authorization concept and an authorization benchmark, for example — the intent decided which paper answered which query, and each linked to the other.

This sounds like search housekeeping, and partly it is. Its larger value was editorial. Forcing every paper to state its single question exposed drafts that were trying to be three papers at once, and splitting them made each one clearer. It also gave readers a natural way to move from concept to test, which is how we hope the library is used.

Lesson 10: Write for the reader who quotes one paragraph

Most readers will not read a paper end to end. They will search, land on a section, and quote a paragraph — increasingly through an AI assistant that summarizes it. We wrote for that reader. Terms are defined at first use. Important paragraphs name their subject explicitly instead of relying on "this" or "the system". Claims that carry evidential weight get one sentence each, with their label attached. And every section that could be quoted on its own begins with its answer.

The practical test was simple: if this paragraph were the only thing someone saw, would it overstate what we know? Paragraphs that failed the test were rewritten until the evidence label travelled with the claim.

What we would do differently

  • Run the independent read earlier. Some security-flagged papers were checked by targeted search rather than full reading before release. A full independent read should come before, not after, the library is assembled.
  • Smaller, linked releases. Releasing everything at once made cross-references coherent, but it concentrated review load. Future releases can be smaller while keeping the same linking discipline.
  • Measure something sooner. A library of designs is valuable for its questions. The next step that matters is running the first protocol and publishing what happens, including if the result weakens our own hypothesis.

Does this apply to AI-assisted writing generally?

Yes. Much of the drafting and checking in a modern publication pipeline can be assisted by AI systems, and ours was no exception. The lessons above are, in effect, the controls that make that assistance safe: fix the evidence class before generating text, constrain citations to a verified registry, turn style rules into checks, and keep a human review step for the judgments automation cannot make. The same principles shape how we think about AI work in our products. The output of an agent should be checked by something other than the agent's own account — a theme that runs through Ethen VerifiedWork, our benchmark design for AI systems that take action.

Limitations

This is a description of one release by one team. It does not measure whether these practices reduce errors compared with alternatives; it reports what we did and what we noticed. Program-level figures in this article come from the library's own release documentation.

References

  1. Pineau, J. et al. (2020). Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program). arXiv:2003.12206. https://arxiv.org/abs/2003.12206
  2. Gebru, T. et al. (2018). Datasheets for Datasets. arXiv:1803.09010. https://arxiv.org/abs/1803.09010
  3. Ethen Research Lab (2026). Research Publications V1. https://upcube.ai/resources/research
  4. Ethen Research Lab (2026). Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action. Benchmark design; not yet run. https://upcube.ai/resources/research/verifiedwork-benchmark