Skip to content

EthenEthenEthen

Why Ethen Uses System Cards and Research Notes Differently

The difference between a system card and a research paper or note comes down to what kind of evidence each carries. A system card reports what one specific system did under a stated method: which build, which conditions, how they were judged, how many passed and what the result does not support. A research note argues, synthesizes existing evidence or proposes a design; it reports no new measurements. Ethen Research Lab publishes both, and keeps them structurally and visually distinct, because each fails in a different way when quoted carelessly. A system card's narrow result gets over-generalized into a broad claim. A research note's argument gets mistaken for a finding. This article explains the two document types, how to read each, and how Ethen's one published system card shows the discipline in practice.

The difference between a system card and a research paper or note comes down to what kind of evidence each carries. A system card reports what one specific system did under a stated method: which build, which conditions, how they were judged, how many passed and what the result does not support. A research note argues, synthesizes existing evidence or proposes a design; it reports no new measurements. Ethen Research Lab publishes both, and keeps them structurally and visually distinct, because each fails in a different way when quoted carelessly. A system card's narrow result gets over-generalized into a broad claim. A research note's argument gets mistaken for a finding. This article explains the two document types, how to read each, and how Ethen's one published system card shows the discipline in practice.

Key takeaways

  • A system card is measured and narrow. One pinned system, a stated method, counted outcomes, explicit limits.
  • A research note is argued and broad. Synthesis, design reasoning or proposal; no new measurements.
  • Both can be honest. They simply carry different kinds of evidence.
  • Each has a typical misreading. Over-generalizing a card; treating a note as a result.
  • Ethen keeps them apart. Different types, different evidence statuses, different page layouts.

What is a system card?

A system card is a document that reports how a specific AI system behaved when tested in a specific way. The idea grows out of a line of work on AI documentation. In 2019, Margaret Mitchell and colleagues proposed model cards: short documents accompanying a trained model that report its evaluation across conditions and state its intended uses and limits. Datasheets for datasets, proposed by Timnit Gebru and colleagues, did the same for data. System cards extend the idea from a single model to a deployed system — model plus surrounding software, policies and safeguards.

What makes a document a system card, in the sense we use, is not its name but its commitments: it identifies exactly which system was tested, states the method, reports all the outcomes including the failures, and says what the result does not support.

What is a research note?

A research note, as Ethen Research Lab uses the term, is a shorter publication that develops an idea: synthesizing existing literature, laying out design reasoning or proposing an approach. Ethen Research Lab has eight research notes. Five are labeled as research syntheses — analysis of existing evidence with no new Ethen measurements — and the others are architecture proposals or a technical note drawing on literature and design decisions. None reports new Ethen measurements.

Research notes are valuable because most important questions start as arguments. A note on Unknown Effects in Autonomous AI Systems, for example, argues that a timeout is not permission to retry — an idea that can shape system design long before anyone measures its effect. But a note is not evidence that the idea works. Figure 1 compares the two types.

Table comparing a system card with a research note on six questions, with type of evidence highlighted: measured under a stated method versus argument and synthesis with no new measurements.
Figure 1. Two documents can be equally honest and carry very different kinds of evidence.

Why the distinction matters

The distinction matters because claims change as they travel. A sentence written carefully in one document is quoted in a summary, paraphrased in a post and compressed by an AI assistant. At each step, qualifiers tend to fall away. The two document types are vulnerable in opposite directions.

A system card is vulnerable to over-generalization. "On this pinned build, 194 of 197 enumerated conditions matched" can become "Ethen's agents pass 98% of safety tests" — a claim about every build, every condition and real-world safety that the card explicitly does not support.

A research note is vulnerable to promotion. "We argue that agents should treat unknown outcomes as unknown" can become "Ethen research shows agents should..." — turning an argument into a finding.

Keeping the types distinct, and labeling each clearly, gives every downstream reader the qualifier they need. We describe the broader principle of keeping research and product claims apart in Why Ethen Keeps Research Separate From Product Claims.

Ethen's system card, as an example

Ethen Research Lab has published one system card: the AgentTrustBench system card. It is worth walking through, because it shows what a narrow, honest result looks like.

The question. Does one pinned Ethen build enforce its stated autonomy boundaries for allow and deny decisions, stop conditions, spend limits and recovery?

The system. One specific build, pinned at a fixed snapshot. Not the product in general, not later builds.

The method. 197 enumerated conditions were run once and then replayed in a different order. A condition passed only if the outcome matched both a prediction recorded before the run and a separate oracle, on both passes. Disagreements were kept rather than discarded. No statistical test was applied.

The counts. 194 of 197 conditions passed. Two conditions disagreed with the prediction and one with the oracle. None were unstable across the two passes, and none errored.

The limits. The card states that the results apply only to that build and that enumerated set; that they do not support a population claim, a product ranking or a live-traffic claim; that it is not a human-trust study; and that the raw records and scorer stay internal.

Notice what the card does not do. It does not round 194 of 197 to a percentage in its headline. It does not call the boundaries "secure". It lists the three non-passing conditions rather than hiding them. And it says plainly what it cannot tell you. That is the standard a system card should meet, and it is also why the result is narrow: honesty and narrowness go together.

How to read a system card

Five questions, shown in Figure 2, get you most of the way to reading any system card correctly — Ethen's or anyone else's.

Five-step reading guide for a system card: which system, what question, what method (highlighted), what counts, and what limits.
Figure 2. A result is only as general as its method and its stated limits allow.

Which system? The exact build or version. Results do not transfer automatically to other versions.

What question? The specific behavior or boundary tested. A card about autonomy boundaries says nothing about answer quality.

What method? How conditions were chosen, run and judged. This is where most of a result's meaning lives. Enumerated conditions designed by the system's builders test what the builders thought to test.

What counts? All the outcomes, including failures and errors. A card that reports only passes is a press release.

What limits? What the card says the result does not support. Read this section first.

How to read a research note

Research notes need a different reading discipline.

Find the evidence status. On every Ethen Research Lab page it sits at the top: research synthesis, proposal, or similar.

Separate the argument from the evidence. Which claims rest on cited external work, and which are the authors' reasoning?

Check the references. A synthesis is only as strong as the sources it synthesizes.

Look for the testable version. A good note points to how its idea could be tested — often a protocol or benchmark design in the same family.

Quote with the qualifier. "Ethen Research Lab argues..." or "A research note proposes..." — not "Ethen research shows".

Common misreadings, and how to correct them

It is easier to see the distinction in examples. Each pair below shows a tempting misreading and an accurate version.

Misreading: "Ethen's agents pass 98% of safety tests." Accurate: "On one pinned build, 194 of 197 enumerated autonomy-boundary conditions matched both a pre-run prediction and a separate oracle." The misreading converts a count into a percentage, a build into "agents", and boundary conditions into "safety tests".

Misreading: "Ethen research shows timeouts should never be retried." Accurate: "An Ethen Research Lab research note argues that an unknown outcome should be reconciled before a consequential action is retried." The misreading turns an argument into a finding and a nuanced rule into an absolute one.

Misreading: "Ethen's benchmark proves its agents handle delegation." Accurate: "Ethen Research Lab has published a benchmark design for evaluating delegation and approval; it has not yet been run." The misreading turns a design into a result.

Misreading: "The system card certifies Ethen as safe." Accurate: "The system card reports one offline enumeration on one build. It is not a certification, a ranking or a human-trust study." The card says this itself; the misreading ignores the limits section.

Questions to ask about any AI system card

System cards from any organization can be read with the same skepticism. Beyond the five reading questions above, a few more help.

Who designed the conditions? Tests designed by the system's builders reflect what the builders anticipated. Independent or adversarial conditions are stronger evidence.

Were failures kept? A card should report every condition run, including those that failed or errored, not a curated subset.

Was anything decided in advance? Predictions or pass criteria recorded before the run make results harder to rationalize afterward.

Can anyone check it? If the raw records stay private, as they do for Ethen's card, the card should say so, and readers should weigh the result accordingly.

What changed since? A card describes one version. Ask whether later versions have been tested.

Where both sit on the evidence scale

Ethen Research Lab labels every publication with an evidence status. Figure 3 places the document types on that scale.

Five bands ordered by strength of evidence: measured results, system evidence (highlighted), research synthesis, protocols and benchmark designs, proposals and hypotheses.
Figure 3. Most research notes sit in the synthesis band and a few are proposals; the system card sits higher, on a very narrow base.

The system card is the only publication labeled system evidence. No publication is labeled as reporting new measured results more broadly. Most research notes are syntheses; some are proposals. Protocols and benchmark designs describe experiments not yet run. Proposals describe untested ideas.

The scale is about kind of evidence, not quality of thinking. An excellent research note can be more useful than a narrow system card. But only the card can tell you what a specific system actually did.

How the pages keep them apart

The Research Lab uses three mechanisms to keep the two types distinct.

Different types and statuses. The system card is typed as a system card and labeled system evidence. Research notes are typed as research notes and labeled with their own status.

Different layouts. The system card has its own page layout, leading with the question, build, method and counts, and marking measured results explicitly. Research notes use the standard publication layout with thesis and abstract. We describe the page design in Inside the Redesign of Ethen's Research Publications.

Different metadata. Scholarly papers and the system card are described differently in the pages' structured data, and neither claims peer review.

How this relates to release certificates

Ethen also publishes a third kind of evidence document outside the Research Lab: release certificates, which record dated, scoped evidence about a specific product release — build gates, boundary checks and what was not tested. We explain them in What a Release Certificate Actually Proves at Ethen. Release certificates share the system card's discipline — specific, scoped, limits stated — but serve operations rather than research. Where content belongs across Blog, Docs and Research is explained in How We Decide Whether Something Belongs in Blog, Docs, or Research.

What we will publish next

As Ethen Research Lab's protocols are run, the archive will need more measured documents: results reports for experiments and, where appropriate, further system cards for specific builds. Each will need to meet the same standard: exact system, stated method, all outcomes, explicit limits. A results report for a protocol will also need to say whether the experiment followed the published protocol and, if not, how it differed. We are not announcing specific publications or dates here.

Tradeoffs and limitations

Narrow results are less exciting. A card that refuses to generalize makes a weaker headline. That is the point.

Builder-designed conditions have limits. Conditions enumerated by the system's own team test what the team anticipated. Independent and adversarial testing would strengthen any card.

Types simplify. Some publications mix synthesis and proposal; each carries its primary type and status.

One card is one card. A single system card on one build is a start, not a track record.

FAQ

What is a system card in AI? A document that reports how a specific AI system behaved under a stated testing method, including all outcomes and the limits of what the result supports.

How is a system card different from a research paper? A system card reports measured behavior of one specific system. A research paper or note may report experiments, synthesize literature or argue for an approach; Ethen's research notes report no new measurements.

What is the difference between a model card and a system card? A model card documents a trained model. A system card documents a whole system — model plus software, policies and safeguards — as tested.

What does a research note mean at Ethen Research Lab? A shorter publication that synthesizes evidence, lays out design reasoning or proposes an approach, with its evidence status stated at the top.

What did Ethen's system card find? On one pinned build, 194 of 197 enumerated autonomy-boundary conditions matched both the pre-run prediction and a separate oracle on two ordered passes. It makes no claim beyond that build and that set of conditions.

References

  1. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model Cards for Model Reporting. Proceedings of FAT\ 2019*. https://doi.org/10.1145/3287560.3287596
  2. Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for Datasets. Communications of the ACM, 64(12), 86–92. https://doi.org/10.1145/3458723
  3. Ethen Research Lab (2026). AgentTrustBench system card: autonomy boundaries on one pinned build. System card. https://upcube.ai/resources/research/agent-trust-boundaries
  4. Ethen Research Lab (2026). Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry. Research note. https://upcube.ai/resources/research/unknown-effects
  5. Ethen Blog. What a Release Certificate Actually Proves at Ethen. https://upcube.ai/blog/what-a-release-certificate-actually-proves-at-ethen