Company
Why Ethen Research Lab Publishes Its Work in PublicEthen Research Lab publishes its work in public so that our claims can be checked, not just read. Every publication states what kind of evidence it contains — a measured result, a research synthesis, a proposal, a protocol or a benchmark design — and the first library of 40 papers says plainly that none of them reports a new measured Ethen result. Publishing that way does four things: it holds our claims to the evidence we actually have, lets others inspect our methods, commits us to how a hypothesis will be tested before any data arrive, and keeps research clearly separate from product claims. We also say what we keep private and why.
Product
How to Explore Ethen Research Lab: Programs, Evidence Labels and Reading PathsThe fastest way to read Ethen Research Lab well is to check two labels before reading anything else: the publication type (position paper, research note, proposal, technical report, methods paper, protocol, benchmark design, survey or system card) and the evidence status (measured result, synthesis, proposal, protocol, or external survey). Together they tell you what kind of claim the paper can make. Then filter the archive by research program to find papers on your topic, and use a reading path to follow a question from concept to benchmark to experiment. This guide explains each label, the programs, how related papers fit together, and where to start for your role.
Engineering
What We Learned Publishing 40 Research Papers at OnceReleasing Ethen Research Lab's first 40 papers together taught us that the hard part of publishing research at volume is not writing — it is keeping every claim matched to its evidence. The lessons that mattered most were procedural. Decide each paper's evidence status before drafting, and let the title follow it. Cite only sources whose claims have been checked against the source itself. Automate the structural checks, but know exactly what automation cannot catch. Treat redaction as its own review pass. Route papers that touch law, privacy or security to the right reviewers. And write down honestly what was not read in full. This retrospective describes that workflow so other teams can borrow it.
Company
Why Ethen Keeps Research Separate From Product ClaimsEthen keeps research and product claims apart because they rest on different kinds of evidence, and readers make different decisions based on them. A product claim says what an Ethen product does today, and should be backed by product evidence. A research claim says what Ethen Research Lab is studying, proposing or has measured, and carries an explicit evidence status — often "proposal" or "protocol; not yet run". When the two blur, research gets read as a shipped feature, and product statements borrow credibility from papers that never tested them. So we publish in three lanes — the Ethen Blog, Ethen Research Lab and company writing — and we follow a few simple rules whenever one lane draws on another.
Product
Why Ethen Shows What It Knows—and What It Doesn'tAI products have a built-in honesty problem: their output sounds equally confident whether it is right, wrong, estimated or invented. Ethen treats that as a design problem, not only a model problem. Across the product, we try to keep four states of knowledge apart — known, estimated, unknown and not checked — and to show each for what it is. When a model fact lacks provenance, Ethen Model Intelligence shows Unknown rather than a guess. When an agent's action times out without confirmation, Ethen's mission system records the outcome as unknown rather than as success or failure. When Ethen Research Lab publishes a proposal, it says the proposal is untested. We do this because people rely on AI outputs more than the outputs deserve when uncertainty is hidden, and because a gap shown honestly is more useful than a gap filled with something plausible.
Product
What “Done” Should Mean for an AI AgentFor an AI agent, "done" should mean that every requirement of the task has been met and that something other than the agent's own report shows it. Precisely: a task is complete when each required obligation is supported by evidence at the level of checking it needs — a passing test, a reconciled record, a confirmed delivery, an approved review — or has been explicitly waived by the person who owns the task. Three refinements make the definition usable. Keep execution success (the steps ran), task success (the outcome was achieved) and business success (it produced value) apart. Treat success as provisional until it can no longer be reversed. And report partial and unknown outcomes as what they are, instead of rounding them up to done.
Company
Inside the Redesign of Ethen’s Research PublicationsThe Ethen Research Lab redesign is built around one decision: every publication tells you what kind of evidence it contains before it asks you to believe anything. At the top of each page, before the title's argument or the abstract, readers see the publication type, an evidence status with a one-line explanation, the research program, the date and the institutional author. Figures carry their own evidence labels. References are numbered and point to persistent identifiers. The archive of 41 publications can be filtered by program, publication type, evidence status and topic, and its legend shows a zero beside "measured" because no publication yet reports new Ethen measurements. This article explains those design decisions, what they are meant to prevent, and what we are still working on.
Company
Why We Made Ethen Research Look More Like a Journal Than a BlogWe chose a journal-like research publication format for Ethen Research Lab because blog conventions make it too easy to confuse an idea with a finding. A blog post has a date, an author and a persuasive headline; it rarely says what kind of document it is or what kind of evidence it holds. Journals solved that problem long ago with typed publications, abstracts, numbered references and stable identifiers. We borrowed those conventions, adapted some — stating evidence status in plain words on every page and treating proposals and protocols as first-class publications — and declined one: we do not imply peer review, because our publications have not been through it. This article explains the reasoning, what each convention is for, and where our format deliberately differs from a journal.
Engineering
Building a Public Research Archive That Can Scale Past 100 PapersGood research archive information architecture comes down to five layers, built from the bottom up: a stable identity for every publication, a classification made of independent facets, explicit relations between related publications, a visible version history, and a presentation layer of curated entry points, hubs and filters. Ethen Research Lab has 41 publications today. That is already more than most readers will browse, and the archive is designed to keep working at well over a hundred. The hard part is not the page count. It is keeping every publication findable, keeping every evidence label accurate as work progresses, and making sure links and citations still work years later. This article explains each layer, the decisions we made, and the mistakes we are designing to avoid.
Engineering
What We Learned Publishing 165 Research VisualsDesigning research figures is hardest when most of the research is not results. Ethen Research Lab's 40 papers carry 165 original visuals: 40 header images, each labeled as containing no data, and 125 figures, each carrying an evidence badge inside the image — proposed architecture, experiment design, proposed measurement framework, qualitative matrix, conceptual diagram and a few others. Not one plots a measurement, because none of the papers reports new Ethen measurements. The central lesson was that a figure's form is itself a claim about evidence: a bar chart says "we measured this", even when it was drawn from opinion. So we labeled evidence inside every image, never drew data we did not have, used one restrained visual system, wrote captions that say what a figure does not show, and treated alt text as content. This article explains those lessons and what we would change.
Company
Why Ethen Uses System Cards and Research Notes DifferentlyThe difference between a system card and a research paper or note comes down to what kind of evidence each carries. A system card reports what one specific system did under a stated method: which build, which conditions, how they were judged, how many passed and what the result does not support. A research note argues, synthesizes existing evidence or proposes a design; it reports no new measurements. Ethen Research Lab publishes both, and keeps them structurally and visually distinct, because each fails in a different way when quoted carelessly. A system card's narrow result gets over-generalized into a broad claim. A research note's argument gets mistaken for a finding. This article explains the two document types, how to read each, and how Ethen's one published system card shows the discipline in practice.
Company
How We Decide Whether Something Belongs in Blog, Docs, or ResearchWe decide between blog, documentation and research by asking what the reader came to do. If they want to use something — set it up, call an API, complete a task — it belongs in Docs. If they want structured facts about a model, it belongs in the Model Library. If they want to know what was found or proposed, it belongs in the Research Lab, typed and labeled with its evidence status. If they want to understand what Ethen is building and how it works, it belongs on the Blog. And if they want to know who Ethen is and why it makes the choices it does, it belongs in Company publishing. Each surface has its own evidence standard and freshness rule, and one topic can appear on several surfaces as long as each page owns one question. This article explains the rules and why they matter.
Product
What Ethen Is Doing to Make AI Outputs Easier to VerifyThe practical answer to how to verify AI output is to check claims, not paragraphs: split an answer into individual statements, open each source, find the exact sentence that supports each statement, make sure the support comes from independent sources, and keep anything you could not confirm labeled as unconfirmed. That works, but it is slow, which is why most AI output goes unchecked. Ethen's approach is to make verification cheaper by building it into the output. Research reports attach a status and the supporting passages to each claim, and say plainly what the check cannot do. Model facts without complete provenance show "Unknown" instead of a guess. Agent work is marked complete only after a separate verifier checks it. Verification capacity is reserved before work is spent. And releases come with scoped evidence. This article explains each mechanism, its limits, and what you should still check yourself.
Company
What Ethen Research Lab Is Exploring Beyond AI ProductsEthen Research Lab research areas reach beyond any single product. The Lab is organized into eight published programs — Evaluation and Verification; Trust and Accountable AI Work; Adaptive Intelligence; Context, Skills and Transfer; Model Intelligence and Faros; Data and Learning Systems; Enterprise and Sovereign AI; and AI Security — connected by one thread: how AI systems can learn from work that can be checked. Some questions feed directly into products. Others look further out: how to measure whether automated graders can be trusted, whether skills survive when the underlying model changes, whether organizations can improve AI without exporting their data, whether agents can predict the effects of their actions, and whether automated research can keep hypotheses separate from confirmed findings. Almost all of this work is published as proposals, protocols and benchmark designs; very little is results. This article describes the questions honestly, without claims the evidence cannot support.
Company
Why Not Everything Ethen Researches Needs to Become a ProductThe relationship between research and product development at Ethen comes down to one rule: every research question ends in one of three outcomes — build, publish only, or stop — and only one of those is a product. Research becomes a product when five things line up: evidence from results rather than proposals, a real need from people doing real work, a cost of ownership we can sustain as models and data change, clearance on safety, privacy and rights, and a natural place in an existing product. Much valuable research meets some of those and not others. It may produce a method others can reuse, a benchmark, a safeguard inside Ethen, a design principle, or a negative result that saves everyone time. That is why a research publication from Ethen is never a product announcement, and why "publish only" and "stop" are normal outcomes rather than failures. This article explains the rule and how to read Ethen research with it in mind.
Company
Building a Research Lab as a Small Team in the Age of AI AgentsA small AI research lab can now do work that once needed a much larger team, because AI agents can search and summarize literature, draft and restructure documents, check consistency and citations, produce figures and scaffold experiments. That is how Ethen Research Lab operates. But agent output is not evidence, and a lab that produces more than it can check produces volume, not knowledge. So the operating model rests on three commitments. People own the questions, the judgment about evidence, the methods fixed before results, and the decision to publish. Agents do bounded labor inside those decisions. And the safeguards — a registry of checked sources, automated checks, review by someone who did not produce the work, protocols published before experiments, and evidence labels on every publication — have to scale as fast as the output does. This article explains that model, where agents help, where they mislead, and what we would advise another small lab.
Engineering
When an Agent Action's Outcome Is UnknownAn agent action is interrupted mid-flight: did it happen? Ethen's mission reconciler refuses to guess — it retains the effect as unknown until evidence resolves it.
Models & Intelligence
How to Read an AI Model ComparisonMost model comparisons answer a narrower question than their headline suggests. Here is how to find the real question, check whether the numbers are comparable, and read the gaps.
Company
What We're Building Across Ethen: October 2026 UpdateAs of October 2026, Ethen's work falls into five areas. We are organizing Ethen into focused apps that share one foundation. We are hardening that foundation for AI work that runs for minutes or hours: durable jobs, completion that depends on evidence, honest handling of unknown outcomes, and approvals tied to specific actions. We are building a model knowledge layer so people can choose models with sources rather than guesses. We are extending Ethen to local, on-device AI through Desktop. And Ethen Research Lab now publishes its research in public, with every paper labeled by evidence status. This update separates what our public posts describe as implemented from what is design direction, and it makes no launch or date commitments.
Product
Ethen Code: What We're Building NextEthen Code is moving from assistive coding — explaining code, writing functions, fixing snippets — toward software tasks that end with evidence a person can check. The direction has seven parts: reproduce a problem before changing anything; plan a minimal, reviewable change before executing it; work in isolated environments with scoped permissions; treat tests, builds and review as evidence rather than as a finish line; put approvals in front of merges, deployments and other consequential steps; recover from failures in long-running work without repeating effects; and keep the same task model across Chat, the cloud workspace and Desktop. This article describes that direction. It is not a release schedule, and it makes no availability or date claims.
Models & Intelligence
Why Ethen Is Investing in Model IntelligenceEthen invests in model intelligence because every decision about which AI model to use — made by a person, by Ethen's AI Gateway, or by Ethen's automatic model choice — is only as good as the facts behind it. Model intelligence is the knowledge needed to make that decision: what a model can do, what it costs, how it performs on which kinds of task, and where each of those facts came from. The model landscape changes too quickly, and public comparisons hide too much, for that knowledge to be assembled ad hoc. So Ethen builds it deliberately: every fact with a source and an owner, Unknown shown instead of guesses, eligibility decided before preference, and a long-term aim of connecting model choices to whether the resulting work actually succeeded.
Engineering
What We're Improving About Reliability Across EthenReliability for AI products means more than staying up. For AI work that acts — editing files, sending messages, generating paid media, running for hours — a reliable system must do four things: be available, make sure each real-world effect happens once rather than twice, make sure tasks actually succeed rather than merely report success, and tell people the truth about what is pending, failed or unknown. Ethen's reliability program works on all four layers. Its main parts are safe retries that never duplicate effects, durable jobs that survive crashes and resume cleanly, deliberate fault testing, verification that is budgeted before work begins, release evidence that states exactly what was checked, careful handling of model and provider changes, and honest status everywhere. This article describes that program at a high level and links to the engineering posts behind each part.
Product
Why Long-Running AI Work Needs a Different UX Than ChatLong-running AI work needs a different user experience than chat because chat is built on four assumptions that stop being true once work lasts longer than a few minutes: that the work takes seconds, that you are watching, that the result is a single reply, and that you can judge that reply on the spot. AI agents that research, code, operate tools or carry out multi-step tasks now routinely run for minutes or hours, often while the person who started them does something else. That work needs a job with its own identity and state, progress shown as phases rather than a spinner, decision points that reach you at the right moment with the context to answer, pause, resume and recovery that never repeat an action twice, and a result delivered with evidence for review — not a confident final message.
Engineering
Why Ethen Is Building for Recoverable AI WorkEthen is building for recoverable AI work because failure is a normal part of long-running work, and what happens after a failure decides whether it becomes a brief pause or an incident. Recoverable AI work means that after any interruption — a crash, a timeout, a provider outage, a revoked permission, a person who needs to think — the work is in a known state, and the next move is a deliberate choice: continue from the last good point, reconcile an uncertain outcome, undo a partial effect where a real undo exists, escalate to a person, or stop safely with an accurate account of what happened. Five foundations make that possible: durable state, actions tied to their intent so repeats do not repeat effects, reconciling before retrying, compensation used only where it is meaningful, and clean stopping and hand-over.
Engineering
What Happens When an AI Task Fails Halfway Through?When an AI task fails halfway through, what happens next depends on one question: did anything change in the world before it failed? A well-designed system handles that question in a fixed order. It stops making new changes. It works out what finished, what did not, and what is uncertain. It checks the uncertain steps by asking the system of record what actually happened. Then it chooses the next move — resume from the last good point, retry only where that is safe, finish just the remaining items, undo a partial change where a real undo exists, or stop and ask you — and it tells you only when a decision genuinely needs you. What it should never do is guess, start over from the beginning, or report the task as either finished or failed when it does not know which.
Company
Why Better AI Models Don’t Eliminate the Need for Better ProductsBetter AI models do not make AI products obsolete; they change which parts of a product matter. Each model release removes some reasons to build: prompt tricks that compensated for a weaker model, savings from routing around price differences, interfaces that do little more than pass text to a model. At the same time, better models are trusted with longer and more consequential work, switched more often, and harder to check by eye — which makes the parts of a product that surround the model more valuable. Those parts are knowing what the AI actually did and whether it worked, controlling what it may touch and spend, recovering when long work is interrupted, testing that a model change will not break real work, keeping memory and data under the user's control, and interfaces built around the work rather than the chat. That is where Ethen invests.
Models & Intelligence
What We Look for Before Adding a New Model to EthenBefore adopting a new AI model, Ethen evaluates it against ten questions. Do we know exactly which model and version it is? Does it fill a gap — a medium, a task, a price or speed point, a local or open option — that Ethen cannot already fill well? Does it perform well on tasks like the ones people actually bring to Ethen, not only on public benchmarks? Is it reliable under real conditions? Will it change without notice? Do its data-handling terms and license allow the uses our users need? How does it behave with risky requests and untrusted content? Is its pricing clear and dated? Does it fit how Ethen calls models and tools? And where, if anywhere, should people see it? A model can be listed in our catalog with sourced facts long before it is qualified for a particular use, and qualified long before it is surfaced in a curated place like Ethen Chat. Each step requires its own evidence.
Models & Intelligence
Why Ethen Sometimes Won’t Use the Newest ModelShould you upgrade to the newest AI model? Not automatically. Ethen sometimes holds back from the newest model because a model that is better on average can still be worse on the specific work people rely on — and averages hide exactly those regressions. A new version may follow formats differently, refuse different requests, call tools differently, cost more per finished task, run slower, or come with different terms. Prompts and skills tuned for the previous model may perform worse until they are adapted. So Ethen treats a new model as a candidate, not an upgrade: we test it on representative work task by task, look past the average to regressions, cost and behavior, switch kind of work by kind of work where it wins, keep tested versions pinned and a fallback available, and keep watching after a switch. Sometimes that means adopting a new model within days. Sometimes it means not adopting it at all.
Models & Intelligence
What Makes an AI Model Useful Beyond BenchmarksThe gap between AI benchmarks and real-world performance comes from what benchmarks leave out. A typical benchmark score measures accuracy on a fixed, public set of tasks, under one method, often from a single attempt, without cost or speed. Real work depends on much more: whether the model fits your inputs, outputs and tools; whether it succeeds every time rather than once; whether it is fast enough to work with; whether it actually uses the long context it accepts; what each good result costs; whether its behavior stays stable; whether it stops honestly when it cannot do something; how it handles instructions hidden in content; and whether its terms and deployment options fit your data. Benchmarks are a useful starting point. Usefulness is decided by these other dimensions, most of which you can test cheaply on your own work.
Product
Why Ethen Founder Is About Outcomes, Not More ChatEthen Founder is built around a simple idea: founders do not need another place to talk to AI; they need work carried through to a result they can check. So Founder's unit of work is the job, not the conversation. You state a goal and its limits. Founder plans the steps, does the research and the work, stops for your approval before anything consequential, checks the result against the goal, and hands back an outcome with its evidence, spend and history attached. That design makes Founder a separate app from Ethen Chat rather than a chat mode. It also comes with explicit limits on what Founder may do on your behalf. This article explains the thinking, what is published about Founder's design, and what has not yet been proven.
Company
What It Means to Build an AI Company With AIAn AI-native company is one that uses AI agents throughout its own work — engineering, research, writing, analysis and operations — rather than only selling AI to others. Ethen is built that way. But using agents everywhere is the easy part. The hard part is keeping the company honest while it does: making sure people stay accountable for every decision, agents work within bounded authority, the context they rely on is written down where both people and agents can read it, claims carry evidence rather than confidence, output and spend are measured rather than assumed, and the judgment that decides what is right is hired for, practiced and protected. Those six principles are the same ones we design into Ethen's products, applied to ourselves. This article explains them, where AI helps and misleads inside a company, and the risks we watch for.
Company
Why We’re Choosing Depth Over Feature CountWhen we weigh product depth vs feature count, we choose depth. Feature count measures inventory: how many buttons, modes and models a product can list. Depth measures whether work actually finishes: whether a workflow covers the real task, works reliably on supported inputs, recovers when a step fails, returns a result you can check, keeps its context between sessions, and lets you control what it touches. AI products are unusually prone to feature sprawl, because every new model capability can become a new button in an afternoon. Ethen deliberately does fewer things per surface and tries to do each one completely. We give each kind of work one home, keep Ethen Chat focused, qualify creative workflows before exposing them, and ask a short set of public questions before adding anything. This article explains why, what we mean by depth, and what the choice costs.
Company
The Road to Ethen V5Ethen V5 is the next generation of Ethen, and it is a rethinking of how Ethen is organized and what it promises rather than a visual redesign. The direction has seven threads that move together: specialized products for different kinds of work; a shared foundation for trustworthy, reliable AI work; model intelligence that helps people choose models with sources; research that informs products through labeled evidence; creative work organized as workflows; autonomous work delegated as jobs with approvals and checked outcomes; and long-term coherence across every surface. This is not a dated Ethen V5 roadmap. We do not publish milestone schedules, because each step is gated by its own evidence rather than a calendar, and each product is announced only when its release evidence exists. This article explains the direction, the principles behind it, and how to follow it.