Skip to content

EthenEthenEthen

Building a Public Research Archive That Can Scale Past 100 Papers

Good research archive information architecture comes down to five layers, built from the bottom up: a stable identity for every publication, a classification made of independent facets, explicit relations between related publications, a visible version history, and a presentation layer of curated entry points, hubs and filters. Ethen Research Lab has 41 publications today. That is already more than most readers will browse, and the archive is designed to keep working at well over a hundred. The hard part is not the page count. It is keeping every publication findable, keeping every evidence label accurate as work progresses, and making sure links and citations still work years later. This article explains each layer, the decisions we made, and the mistakes we are designing to avoid.

Good research archive information architecture comes down to five layers, built from the bottom up: a stable identity for every publication, a classification made of independent facets, explicit relations between related publications, a visible version history, and a presentation layer of curated entry points, hubs and filters. Ethen Research Lab has 41 publications today. That is already more than most readers will browse, and the archive is designed to keep working at well over a hundred. The hard part is not the page count. It is keeping every publication findable, keeping every evidence label accurate as work progresses, and making sure links and citations still work years later. This article explains each layer, the decisions we made, and the mistakes we are designing to avoid.

Key takeaways

  • Identity first. Each publication gets one address that never changes or gets reused.
  • Facets, not folders. Program, type, evidence status and topic are independent and combine freely.
  • Relations are explicit. A proposal links to the protocol that would test it and the benchmark that would measure it.
  • Nothing changes silently. Updates, status changes, supersession and withdrawal all leave visible traces.
  • Curate the front door. Featured publications and program hubs matter more as the archive grows.
  • Automate the checks. Consistency across a hundred papers cannot rely on memory.

Why archives break as they grow

Small archives work almost regardless of design. With ten publications, a reverse-chronological list is fine. Problems appear gradually, and by the time they are obvious they are expensive to fix.

Findability collapses. A date-ordered list stops working once readers cannot scan it. People arriving with a question cannot find the relevant publication, and older work disappears below the fold.

Labels drift. Evidence statuses that were accurate at publication become wrong as protocols are run and proposals are tested — unless something forces them to be updated.

Links rot. Renamed pages, reorganized sections and changed address schemes break inbound links and citations. Every broken link damages trust in the archive.

Duplicates multiply. Without clear relations, closely related publications compete with each other in search and confuse readers about which is current.

Silent edits erode trust. If a publication changes without a visible note, anyone who cited the earlier version is left with a claim the page no longer makes.

The five layers in Figure 1 are each a defense against one or more of these problems.

Five stacked layers: identity (highlighted), classification, relations, versions and presentation.
Figure 1. Get the bottom layer wrong and every layer above it breaks as the archive grows.

Layer 1: identity

Every publication needs one permanent address. Tim Berners-Lee made the case for this in a short 1998 note titled "Cool URIs don't change": addresses should be designed so they can last for decades, which means leaving out anything likely to change, such as file formats, status words or the current organizational structure.

For Ethen Research Lab, that principle translates into a few rules.

Addresses are short, descriptive slugs. Each publication lives at one address under the research section, derived from its subject rather than its status or program.

Status is never in the address. A protocol that later reports results keeps its address. Putting "proposal" or "draft" in an address would force a change exactly when the publication becomes most worth linking to.

Programs and topics are not in the address. Research programs may be reorganized. Topics may be renamed. Neither should break a link.

Addresses are never reused. If a publication is withdrawn, its address keeps explaining that, rather than being given to something new.

Identity is the cheapest layer to get right at the start and the most expensive to fix later, which is why it comes first.

Layer 2: classification with independent facets

The natural instinct when organizing documents is a hierarchy of folders: programs, then topics, then types. Hierarchies break quickly, because real publications belong in several places at once. A benchmark design about agent memory belongs under evaluation, under memory and context, and under benchmark designs simultaneously.

The alternative is faceted classification, where each publication is described by several independent properties that readers can combine. Research on faceted navigation, notably work by Marti Hearst and colleagues on browsing large collections, found that exposing several independent facets lets people start from whichever dimension matches their question and then narrow by any other. Ethen Research Lab uses four facets, shown in Figure 2.

Table of four independent facets with the question each answers and examples: research program, publication type, evidence status (highlighted) and topic.
Figure 2. Facets that answer different questions can be combined without creating a tangled hierarchy.

Research program answers which line of work a publication belongs to. Publication type answers what kind of document it is. Evidence status answers how strong its evidence is. Topic answers what it is about, using the same topic vocabulary as the rest of the Ethen site so research and blog content connect.

The key design rule is that each facet answers a different question. When two facets overlap — say, a topic that is really a program — readers get confused about which to use and editors classify inconsistently. Keeping them independent makes both browsing and classification simpler.

Evidence status as a facet deserves special mention. Most archives classify by subject. Classifying by strength of evidence lets readers ask questions that matter for trust: show me everything with results; show me every planned experiment; show me what is still only a proposal. We describe how this appears on the page in Inside the Redesign of Ethen's Research Publications.

Layer 3: explicit relations

Research comes in families. A position paper argues that something matters. A protocol describes how to test it. A benchmark design describes how to measure it. Eventually, results report what happened. If these publications are not explicitly linked, readers find one and miss the others, and search engines treat them as competitors.

The archive handles this with explicit relations rather than relying on similarity. Each publication lists its related research, and the relation types are meaningful: this protocol tests that proposal; this benchmark design measures that capability; these results answer that protocol. Ethen Research Lab's 40-paper release was organized this way from the start — each paper answers one question and belongs to a family — which we describe in What We Learned Publishing 40 Research Papers at Once.

Explicit relations also solve a search problem. When several publications address related questions, each should own one specific question and link to the others for the rest. That keeps the archive from competing with itself.

Layer 4: versions and visible change

Research publications change. Errors are corrected. References are added. Protocols are run and gain results. Proposals are superseded by better ones. Occasionally, something should be withdrawn. Figure 3 shows the life of a publication record as we design for it.

Five-stage lifecycle of a publication record: published, updated with dated notes, status changes (highlighted), superseded with a pointer, and withdrawn with an explanation.
Figure 3. Nothing disappears silently; every change leaves a visible trace.

The principle is that every change leaves a visible trace.

Updates carry dated notes. A correction or addition is noted on the page with its date and what changed.

Status changes are prominent. When a protocol gains results, its evidence status changes, and the change is visible both on the page and in the archive's filters. This is the most important kind of update, because it changes how much the publication should be trusted.

Supersession points forward. If a newer publication replaces an older one, the older page stays at its address and points clearly to the newer one.

Withdrawal explains itself. If a publication is withdrawn, its address remains, with an explanation of why.

Preprint servers such as arXiv have long handled this with numbered versions, so that a citation can point to the exact text that was read. Version history and version-specific citation are on our list of improvements for the Research Lab; the principle of visible change applies already.

Layer 5: presentation that scales

The top layer is what readers actually see. As the archive grows, it has to do more curation, not less.

Featured publications. A small, deliberately chosen set of starting points for new readers. With forty publications, a reader can browse. With a hundred and fifty, they need someone to say where to start.

Program hubs. Each research program can have its own page summarizing its question, its publications and their evidence status. Hubs become essential once a program has more than a handful of publications.

Filters. The four facets, available on the archive page, with clear counts.

Machine-readable indexes. Search engines and AI systems discover publications through sitemaps and structured data. Those should describe exactly what the page shows — no more. Ethen's research pages mark scholarly papers with a type that describes scholarly writing without asserting peer review, and do not encode evidence status in properties that would misrepresent it.

Reader guides. A guide to reading the archive, like How to Explore Ethen Research Lab, becomes more valuable as the archive grows.

Automated checks that scale

A hundred publications cannot be kept consistent by memory. The checks that matter most are the ones that can be stated precisely and run on every change.

Every publication has a type, an evidence status, a program and at least one topic. Missing classification is the most common cause of findability problems.

Every internal link resolves. Broken links between publications are caught before they reach readers.

Every reference has an identifier where one exists. References without DOIs or arXiv identifiers are flagged for review.

Every figure has alternative text and an evidence label. Accessibility and honesty both depend on it.

Evidence status matches content. This is the hardest check to automate fully. Simple rules help — a publication labeled as reporting results must contain a results section — and editorial review handles the rest.

We describe the editorial version of this discipline in What We Learned Publishing 40 Research Papers at Once.

What changes between 40 and 150 publications

It helps to picture the archive at three sizes. The thresholds are approximate and illustrative, but the shifts are real.

Around 40 publications — where Ethen Research Lab is today — a motivated reader can still scan the whole list. Filters are a convenience. A single landing page with a featured set works well.

Around 80 publications, scanning stops working. Most readers will use filters or arrive from search. Evidence-status counts become one of the most useful things on the page, because they summarize the state of the program in one glance. Families of related publications need explicit links, or readers will find a proposal without ever seeing the protocol that tests it.

Around 150 publications, the archive needs a second level of navigation. Program hubs carry more of the load than the main list. Older publications that have been superseded need to point clearly to their successors, and reader guides become essential rather than optional. Editorial consistency across that many documents is only possible with automated checks.

Designing the five layers now, while the archive is small, is far cheaper than retrofitting them later. Every address created today is one that will need to keep working when the archive is three or four times larger.

Search engines and AI assistants are now among the most important readers of a research archive, and the five layers help them too.

Stable addresses accumulate links and citations over time instead of losing them to redirects. Explicit relations and one-question-per-publication reduce the chance that several pages compete for the same query. Facet pages and program hubs give search engines clear topical structure. Modest, accurate structured data means summaries generated from the metadata do not overstate the evidence. And visible version notes help both people and machines understand which statement is current.

The goal is not to rank for as many queries as possible. It is for each publication to be the best answer to one specific question, and to be found by people asking it.

Mistakes we are designing to avoid

Several common mistakes shaped these decisions.

Encoding status in addresses or titles. "Draft", "preliminary" or "v2" in an address guarantees a broken link later.

Deep hierarchies. Nested folders force each publication into one place and make reorganization painful.

Tag sprawl. Unlimited free-form tags create dozens of near-duplicates. A small, controlled vocabulary per facet works better.

Similarity-only related links. Automatically generated "related" lists are noisy. Explicit relations carry meaning.

Silent corrections. Changing a claim without a note is the fastest way to lose a careful reader's trust.

Letting the front page become a list. At scale, the archive's entry page needs curation, not just a feed.

Tradeoffs and limitations

Classification takes editorial time. Four facets per publication, plus relations, is more work than tagging.

Controlled vocabularies need maintenance. Programs and topics evolve; changing them without breaking filters takes care.

Some features are still to come. Version history and version-specific citation are planned improvements, not finished features.

Automation cannot judge evidence fully. Checks catch missing labels and broken links; whether a label is right still needs human review.

FAQ

How should a company organize its research papers online? Give each paper a permanent address, classify it with a few independent facets such as program, type, evidence status and topic, link related papers explicitly, make every change visible, and curate entry points for new readers.

How do you version research publications? Keep the same address, add dated update notes, make status changes prominent, point superseded publications to their replacements, and keep withdrawn publications at their address with an explanation.

What is faceted navigation? A way of browsing in which each item has several independent properties, and readers can start from any property and narrow by the others.

Why classify research by evidence status? Because readers need to know how much to trust a publication, and filtering by evidence lets them find, for example, everything with results or every planned experiment.

How many papers does Ethen Research Lab have? 41 publications as of October 2026: 40 papers and one system card.

References

  1. Berners-Lee, T. (1998). Cool URIs don't change. W3C. https://www.w3.org/Provider/Style/URI
  2. Yee, K.-P., Swearingen, K., Li, K., & Hearst, M. (2003). Faceted metadata for image search and browsing. Proceedings of CHI 2003. https://doi.org/10.1145/642611.642681
  3. arXiv. Replacing an article (versioning policy). https://info.arxiv.org/help/replace.html
  4. Schema.org. ScholarlyArticle. https://schema.org/ScholarlyArticle
  5. Ethen Research Lab. Research Lab index. https://upcube.ai/resources/research