Skip to content

EthenEthenEthen

Showing Unknowns in Ethen Model Intelligence

When a benchmark cell has no defensible value, Ethen Model Intelligence says so. This is how provenance and missing results control what readers see.

When a benchmark cell has no defensible value, Ethen Model Intelligence says so. This is how provenance and missing results control what readers see.

A benchmark table with gaps can look unfinished. In Ethen Model Intelligence, the gaps are deliberate. If a value arrives without a source URL, without a retrieval date, without a methodology statement, or with low confidence, the display layer does not interpolate, carry forward an old score, or borrow a number from a nearby benchmark. It renders Unknown and keeps the reason attached to the cell. That behavior comes from a small set of explicit gates in the provenance and evaluation code, plus a candid evidence record for the current milestone that says which inputs are actually present.

This article walks through those gates: what counts as known, what provenance carries, why chart-derived values often stay unknown, how comparability and ranking stay narrow, and why missing results produce empty states rather than silent orderings. It also states the limits plainly. There is no original Ethen index in these results, and matching null methodology fields are not evidence of complete comparability.

Known is a conjunction, not a default

The central display primitive is an evidence state with two outcomes: known or unknown. Each state carries the underlying value or null, a display string, an optional reason, and a provenance record. The constructor that produces this state takes a candidate value, a display string, a provenance object, and an optional reason, then checks five conditions together.

To be known, the value must be non-null, the provenance must include a source URL, the retrieval timestamp must parse as a valid date, the methodology field must be non-empty, and the confidence must not be low or unknown. If any one of those fails, the constructor returns unknown, discards the candidate value by setting it to null, sets the display string to Unknown, and preserves a reason. The default reason is direct: a sourced value, retrieval date, and methodology are required.

That conjunction matters for readers. A cell can have a plausible number in an upstream file and still show Unknown because the retrieval date is missing or unparseable. A cell can have a source link and a number but still show Unknown because the methodology string is empty. A cell can have all three and still show Unknown because confidence is low or unknown. The display does not rank these failures or pick the closest pass. Any missing conjunct forces the same visible outcome, with the reason field as the only differentiation.

For engineers, this is a fix list: supply all five inputs or the cell stays Unknown. No strong field compensates for a missing one.

What provenance actually records

Each provenance record carries source URL, source label, retrieval timestamp, methodology, and confidence, with optional fields for source type, effective and expiry timestamps, a source hash, and a methodology version. The required fields drive the known-or-unknown decision above. The optional fields add lifecycle and identity information without changing that gate.

Source URL and retrieval timestamp together answer where and when. A named source without a retrievable URL does not pass. A URL without a parseable retrieval date does not pass. Methodology answers how the number was produced, in free text. Confidence summarizes how much the pipeline trusts the combination, on a four-level scale of high, medium, low, or unknown, where only high and medium can support a known display.

The optional fields — source type, effective and expiry timestamps, source hash, methodology version — add lifecycle detail but are not required by the gate and are not guaranteed on every record. A known cell means the five required inputs passed, not that versioning or expiry metadata is complete. Comparing two known cells across time still requires checking methodology versions, source types, and freshness in the evaluation layer.

Why chart-derived values often stay unknown

A separate helper builds provenance for chart-derived data. It takes a canonical URL, a retrieval timestamp, a source type, a methodology string, and a description. It maps the source type to a label: a JSON-LD dataset source becomes Published benchmark dataset, while anything else becomes Committed normalized dataset. It validates the retrieval timestamp and falls back to null when parsing fails. It takes methodology from the methodology input or, if absent, from the description input. It sets confidence to medium for JSON-LD dataset sources and to unknown for everything else.

The consequence is immediate. Any chart-derived value that flows through this helper without the JSON-LD dataset source type arrives with unknown confidence, which fails the known-or-unknown gate regardless of the other fields. The display then shows Unknown even if a URL, date, and methodology-like description are present. That is not a rendering bug. It is the pipeline refusing to upgrade a committed normalized dataset into displayed evidence on description text alone.

This explains why a benchmark with a visible methodology paragraph can still show Unknown: if that text arrived as a description fallback and the source type yields unknown confidence, the gate fails on confidence. The fix is a qualifying source type and evidence trail, not a rewritten paragraph.

The registry currently describes one benchmark

The evaluation module defines benchmark definitions, benchmark results, and ranking helpers under explicit schema and normalization policy versions. A benchmark definition carries an identifier, name, task, description, unit, scale, direction, methodology, methodology version, source type, and comparability group. A benchmark result carries an identifier, model reference, benchmark reference, numeric score, display value, effective and recorded timestamps, methodology version, evidence state, and source type. Ranked evaluations add a nullable rank and an eligibility flag.

The current registry contains exactly one definition: the Artificial Analysis Intelligence Index, categorized as a reasoning task, measured in index units, with higher-is-better direction, described as a composite external benchmark as described by the committed dataset, with methodology stated as an independent test run by Artificial Analysis on dedicated hardware, sourced as an external benchmark, and grouped under its own comparability group. Its methodology version is null. Its scale is null.

The milestone evidence record confirms the narrow scope: benchmark count one, task categories limited to reasoning, and a data-gaps note that only one benchmark has identity, methodology, and comparability evidence represented in the current registry. The result-count field does not give a number. It states that no canonical chart or result source with model, benchmark, score, and evidence linkage is present, and that the raw-score architecture validates supplied results without silent ranking.

For display, this means no second benchmark to compare against, no Ethen-authored composite index in these results, and no canonical result set to fill a leaderboard. Comparison, ranking, and normalization paths exist, but the joined data they would order is not present as a canonical source.

Comparability requires same test, same version, same source type, fresh on both sides

Two results are comparable only when five conditions hold together: both reference the same benchmark identifier as the definition under consideration, the definition direction is known rather than unknown, both carry the same methodology version, both carry the same source type, and both have fresh evidence state. The helper that checks this takes the definition and two results and returns a boolean. There is no fuzzy matching on benchmark names, no cross-benchmark equivalence, and no grace period for stale data.

Each condition blocks a misreading: same-benchmark matching blocks cross-test mixing, known direction blocks ordering when better is undefined, version equality blocks mixing procedures, source-type equality blocks mixing external, provider-reported, Ethen-internal, and gateway-runtime collection conditions, and fresh-on-both-sides blocks stale, unknown, or invalid inputs.

The version check needs care because the registry entry carries a null methodology version. In code, null equals null, so two null-version results satisfy that conjunct without proving they followed the same procedure. Null methodology equality is not evidence of complete comparability: matching nulls mean version information is absent, not that versions match.

Source-type equality is similarly strict. The registry entry is typed as an external benchmark. A result typed as provider-reported, Ethen-internal, or gateway-runtime cannot be comparable to it under this definition, even if the benchmark identifier and methodology version match and both sides are fresh. The pipeline treats source type as part of the measurement conditions, not as administrative metadata.

Ranking eligibility is narrower than having a score

Ranking builds on comparability but adds its own filter. For a given definition and a list of results, the ranking function first excludes everything when direction is unknown. Otherwise it keeps only results that reference the definition identifier, carry fresh evidence state, match the definition source type, match the definition methodology version, and have a finite numeric score. It sorts the survivors by direction, assigns dense rank positions, and returns the original list annotated with eligibility flags and nullable ranks. Ineligible results stay in the output with eligibility false and rank null. They are not dropped, reordered, or silently promoted.

The leaderboard accessor wraps this logic by looking up the benchmark identifier in the registry and delegating to the ranking function, returning an empty list when the identifier is not found. The milestone evidence record adds an explicit user-interface expectation: the leaderboard projection calls this canonical function and renders an empty state when canonical results are unavailable. Stale and unknown inputs are excluded from ranking rather than demoted to the bottom.

Consider what this means for three plausible inputs. A result with a finite score but stale evidence state is ineligible and keeps a null rank. A result with a finite score and fresh state but a provider-reported source type, against the external-benchmark definition, is ineligible and keeps a null rank. A result with a finite score, fresh state, and matching source type but a non-matching methodology version is ineligible and keeps a null rank. In each case the score may be visible elsewhere as a raw value with its own provenance, but it does not participate in ordering. The leaderboard shows only the survivors, or nothing when there are no survivors.

This prevents a sorted table from implying a definitive order from incomparable inputs: ineligible rows stay annotated, and missing inputs yield an empty state rather than an invented order.

No cross-benchmark normalization in this milestone

The normalization helper operates strictly within one benchmark. It takes a definition, a target result, and a set of comparable results, and returns null unless direction is known and every member of the set is comparable to the target under the comparability check above. When those preconditions pass, it rescales the target score to the zero-to-one range defined by the minimum and maximum of the set, honoring direction so that higher-is-better and lower-is-better both map best to one. When minimum equals maximum, it returns one rather than dividing by zero.

The module comment is explicit that this milestone deliberately provides no cross-benchmark universal normalization. There is no function that takes a reasoning score and a coding score and produces a single blended intelligence number. The policy version records this stance: normalization policy mi-normalization-r4.0 governs within-benchmark rescaling only.

Two higher-is-better numbers on different tests therefore stay two separate facts with separate provenance and comparability groups. A composite blending them would need its own methodology, versioning, and evidence. The current registry, with one benchmark and no canonical joined source, does not support such a composite.

What missing, stale, and invalid actually do to the page

Results carry an evidence state of fresh, stale, unknown, or invalid. Only fresh participates in comparability and ranking. Stale, unknown, and invalid are excluded from ordering. This is consistent with the provenance gate: a value can exist as a stored number while displaying as Unknown or while sitting out of a leaderboard because its evidence lifecycle does not support comparison.

The milestone evidence record summarizes this as excluded from ranking, and pairs it with the empty-state expectation for the leaderboard projection. The display contract is therefore threefold. Raw values validate without silent ranking. Ineligible values keep null ranks. Missing canonical inputs produce an empty leaderboard rather than a partial or speculative one.

For example, suppose a page lists the Artificial Analysis Intelligence Index with three candidate rows: one fresh external-benchmark row with a matching methodology version, one stale external-benchmark row, and one fresh provider-reported row. Only the first is ranking-eligible. The leaderboard shows that survivor alone, or an empty state with an explanation — never a one-two-three ordering of all three. If the fresh row later lacks provenance at the display layer, it can also show Unknown in value cells while ranking treats eligibility separately.

Missing linkage behaves the same way. When scores exist without the required model, benchmark, score, and evidence linkage — the current canonical situation per the evidence record — the correct display is an empty state saying canonical results are unavailable, with raw-score validation reserved for supplied results that carry the required fields.

Reading a cell without overreading it

Because known and ranking-eligible are separate, readers need two questions for every number they see. First, is the cell known, meaning value, source URL, valid retrieval date, methodology, and sufficient confidence all passed? Second, if the cell sits in a ranked context, is the row eligible, meaning same benchmark, known direction, matching methodology version, matching source type, and fresh on every compared side?

A known cell in an unranked context is a sourced fact with a date and method. It supports statements like this source reported this score on this date under this methodology. It does not by itself support cross-model ordering unless the ranking preconditions also hold. A ranked row implies those preconditions held for the compared set at ranking time, within one benchmark and one source type. It does not imply cross-benchmark standing, and it does not imply that matching null methodology versions captured full procedural identity.

When methodology versions are null on both definition and results, the version-equality check passes vacuously and contributes no discrimination; comparability then rests on the remaining conjuncts plus the external methodology description. Treating that pass as proof procedures matched would overread the code.

The single-benchmark registry further limits claims. A leaderboard under the Artificial Analysis Intelligence Index speaks only to that index as described by the committed dataset and its stated independent test procedure — not to coding, math, knowledge, instruction following, multilingual, vision, tool use, long context, or safety tasks, which exist as type labels but have no registry definitions. The evidence record lists reasoning as the sole task category.

Where these surfaces live

The web route registry defines the addressable surfaces for model intelligence content, including the model intelligence home, models, providers, benchmarks, leaderboards, and categories, each with index and detail patterns, alongside the public model library and documentation routes. Those routes establish where benchmark and leaderboard content can appear. They do not by themselves populate those pages with canonical results. The evidence record governs what the leaderboard projection should do when canonical inputs are absent: call the canonical ranking function and render an empty state.

A route can therefore render correctly while showing Unknown cells or an empty leaderboard: surface existence and result availability are independent facts. The marketing index describes model intelligence as a place to compare capabilities, benchmarks, pricing, and evaluations where available — and availability is determined by the provenance and evaluation gates, not by the route.

Evidence and limits, stated together

The strengths of this design are specific. Provenance failures produce visible Unknowns with reasons instead of silent gaps or invented numbers. Chart-derived data without adequate source typing stays unknown instead of entering displays on description text alone. Comparability requires same benchmark, known direction, same methodology version, same source type, and fresh states, which blocks the most common accidental mixings. Ranking eligibility mirrors those requirements and annotates ineligible rows with null ranks instead of dropping them without trace. Normalization stays within one benchmark and refuses to invent a universal scale. Leaderboards render empty states when canonical results are unavailable, as validated in the milestone record.

The limits are equally specific. Only one benchmark definition is registered, with a null methodology version and a null scale, so version discrimination and unit scaling rest on absent fields. No canonical joined result source with model, benchmark, score, and evidence linkage is present, so leaderboard population has no canonical input to order. Full-repository type checking was pending at the time of the evidence record; only a focused evaluation test had been executed. Confidence assignments for chart-derived data depend on a coarse source-type mapping that yields medium or unknown, which means display eligibility can turn on that single classification. Optional provenance fields for source type, effective and expiry timestamps, source hash, and methodology version are available in the type but not guaranteed on every record.

Two broader cautions follow. First, there is no original Ethen index in these results. The sole registered benchmark is an external composite attributed to its publisher and described procedure, not an Ethen-authored blended score. Any future Ethen index would need its own definition, methodology, versioning, and canonical results before it could be displayed or ranked. Second, null methodology equality across results is not evidence of complete comparability. It satisfies one conjunct in code while leaving procedural identity unverified. Readers and downstream writers should repeat that caveat wherever they discuss matching nulls, rather than letting an equality check stand in for a methods comparison.

Unknown is the honest cell

Showing Unknown is not a placeholder for a number that will arrive on its own. It is the output of explicit checks that ask for a value, a source, a date, a method, and sufficient confidence, and then ask again whether freshness, version, and source type support comparison and ranking. When those checks fail, the interface says Unknown, marks rows ineligible, or renders an empty leaderboard, and keeps the reason nearby. When canonical evidence is absent, as the current milestone record states for joined results, the empty state is the correct rendering, not a defect to work around by sorting whatever numbers happen to be at hand.

For Ethen benchmark provenance, that discipline is the story. Provenance decides whether a fact can be shown. Evaluation decides whether shown facts can be compared or ordered. Missing inputs narrow what can be displayed at each step, visibly and with reasons, so readers see the boundary between what is sourced and what is not yet known.