What Ethen Research Lab Is Exploring Beyond AI Products
Ethen Research Lab research areas reach beyond any single product. The Lab is organized into eight published programs — Evaluation and Verification; Trust and Accountable AI Work; Adaptive Intelligence; Context, Skills and Transfer; Model Intelligence and Faros; Data and Learning Systems; Enterprise and Sovereign AI; and AI Security — connected by one thread: how AI systems can learn from work that can be checked. Some questions feed directly into products. Others look further out: how to measure whether automated graders can be trusted, whether skills survive when the underlying model changes, whether organizations can improve AI without exporting their data, whether agents can predict the effects of their actions, and whether automated research can keep hypotheses separate from confirmed findings. Almost all of this work is published as proposals, protocols and benchmark designs; very little is results. This article describes the questions honestly, without claims the evidence cannot support.
Ethen Research Lab research areas reach beyond any single product. The Lab is organized into eight published programs — Evaluation and Verification; Trust and Accountable AI Work; Adaptive Intelligence; Context, Skills and Transfer; Model Intelligence and Faros; Data and Learning Systems; Enterprise and Sovereign AI; and AI Security — connected by one thread: how AI systems can learn from work that can be checked. Some questions feed directly into products. Others look further out: how to measure whether automated graders can be trusted, whether skills survive when the underlying model changes, whether organizations can improve AI without exporting their data, whether agents can predict the effects of their actions, and whether automated research can keep hypotheses separate from confirmed findings. Almost all of this work is published as proposals, protocols and benchmark designs; very little is results. This article describes the questions honestly, without claims the evidence cannot support.
Key takeaways
- Eight programs, one thread. Learning from work that can be proven.
- Some questions are product-adjacent; some are not. Both are worth asking.
- Most work is questions and methods. Results will be published when experiments run.
- Long horizons depend on near ones. Far questions need evidence from nearer questions first.
- No AGI claims. No benchmark certifies general intelligence, and we set no dates.
The eight programs
Figure 1 lists the Lab's published programs and the question each asks.
Evaluation and Verification asks how we can know whether an AI system's work is right — including whether the automated graders used to check it can themselves be trusted.
Trust and Accountable AI Work asks how agents should act, ask for approval, keep records and recover from failure.
Adaptive Intelligence asks whether AI systems can improve from work that has been verified, rather than from any activity at all. Its central position paper is Verified Adaptive Intelligence: Learning From Work That Can Be Proven.
Context, Skills and Transfer asks what survives when models, tools and context change.
Model Intelligence and Faros asks how a system should choose which model to use for a task, and when learned choices beat good rules.
Data and Learning Systems asks which data is worth learning from, and with what rights.
Enterprise and Sovereign AI asks whether organizations can evaluate and improve AI inside their own boundaries.
AI Security asks how agents stay safe under adversarial pressure.
For a guide to navigating the archive, see How to Explore Ethen Research Lab.
Questions by horizon
Some of the Lab's questions could produce useful answers soon; others are years away. Figure 2 arranges them by horizon.
Near: can we trust the checks?
Much of AI evaluation now relies on automated graders, often language models judging other models' work. If those graders are unreliable, every result built on them is suspect. Ethen Research Lab has published a protocol, How Should We Measure the Reliability of LLM Verifiers?, setting out how to measure false positives, false negatives and calibration against expert judgment — not yet run. A methods paper, Evaluating the Evaluators: Reward Integrity for AI Agents, examines how checks can be gamed.
Related near-term questions concern recovery and proof: when an agent should retry, reconcile, escalate or stop, explored in the Recovery Atlas proposal, and what a verifiable record of completed work should contain, in the Work Receipts technical report.
Middle: what survives change?
AI systems change constantly: models are upgraded, tools are updated, context grows and is compressed. Three questions sit here.
Do skills survive model upgrades? The Capability Transfer Ledger methods paper and a related protocol describe how to measure whether an agent's skills hold up when the underlying model changes. Results are gated until experiments run.
Can context be compressed without losing obligations? Agents working on long tasks accumulate context and must summarize it. The Evidence-Preserving Context proposal asks how to compress memory without dropping commitments the agent still owes.
Can organizations improve AI without exporting data? The Private AI Improvement Without Raw Data Export proposal and a related sovereign improvement protocol explore whether learning can happen inside an organization's boundary.
Far: can agents model what their actions will do?
Longer-term questions concern how agents represent the world they act in.
Digital state models. Can an agent predict the likely effects of an action in software — what will change, what might go wrong, how confident it is — and use that to plan? Predictions must never be treated as fact; they are checked against what actually happens.
Compositional skills. Can an agent combine skills learned separately to handle tasks it has never seen, and can that be distinguished from memorized demonstrations?
Credit assignment over long tasks. When a long task succeeds or fails, which decisions, context choices and corrections were responsible? Answering that well would make learning from experience far more efficient.
These questions connect to the Lab's robotics direction, described in From Screen to Physical World: How We Think About Ethen Robotics.
Furthest: can automated research stay honest?
The furthest question is whether AI systems can help conduct research itself — proposing hypotheses, designing experiments, running them and reporting results — while keeping hypotheses clearly separate from confirmed findings.
Others are exploring this actively. Work on an AI "co-scientist" has described multi-agent systems that generate research hypotheses for experimental testing in biomedicine, and the "AI Scientist" project explored language models carrying out machine learning research end to end, from ideas to papers. These are notable efforts. They also illustrate the central risk: a system that writes fluent papers can make hypotheses look like findings.
Our interest is in the discipline around automated research: starting with software experiments where results can be checked; keeping a lab notebook of what was tried; requiring independent confirmation before anything is called a finding; and never treating an external report of a discovery as confirmed by default.
How the programs connect
The programs are not independent tracks. They form a chain that runs from checking work to learning from it.
Evaluation and Verification provides the measurement foundation: without trustworthy checks, nothing downstream can be trusted. Trust and Accountable AI Work produces the records — approvals, evidence, recovery decisions — that make work checkable in the first place. Adaptive Intelligence asks whether systems can learn from the work those records show to be verified. Context, Skills and Transfer asks whether what is learned survives change. Model Intelligence asks how to choose the right model for each step. Data and Learning Systems asks which data may be used and is worth using. Enterprise and Sovereign AI asks whether all of this can happen inside an organization's boundary. And AI Security asks how the whole chain holds up when someone is trying to break it.
A weak link anywhere weakens everything after it. That is why the nearest-horizon questions, about verification and recovery, receive the most attention: they are the foundation for the rest.
Questions we decided not to pursue, for now
Choosing questions also means declining some.
Training frontier models from scratch. Large-scale pretraining may one day be a research instrument, but each scaling run would need to answer a specific scientific question. It is not where the Lab starts.
Benchmarks designed to rank models on leaderboards. Leaderboards are useful, but the Lab's benchmark designs focus on whether systems can be trusted to act, not on producing a single score.
Claims about consciousness or general intelligence. These are interesting questions, but nothing in the Lab's methods would let it answer them, so it does not try.
Research that requires data we lack rights to use. Every program depends on data with clear permissions. Where those do not exist, the question waits.
These choices can change as evidence and capacity change. When they do, the change will be published.
Why research beyond products?
A product company could reasonably ask why it should study questions that may never become products. We have three reasons.
Products need answers that do not exist yet. Whether an agent's work can be trusted depends on whether its verifiers are reliable — an open research question. Whether an upgrade is safe depends on whether skills transfer — another one. Asking these questions publicly, with methods fixed in advance, is how we expect to get answers we can rely on.
Some questions are bigger than any product. How AI systems should learn from experience, and how to keep that learning honest, are questions for the whole field. Contributing methods and negative results is worthwhile even when they do not become features.
Publishing questions disciplines us. A question published with a protocol cannot quietly become a claim. We explain why we publish in public in Why Ethen Research Lab Publishes Its Work in Public.
What happens to a question
Figure 3 shows how a question moves through the Lab.
A question is stated publicly with why it matters. A protocol fixes the method before any results. Results are published whatever they show, including null results. Then comes a decision: build something on the result, publish it as a contribution without building anything, or stop. Earlier publications are updated visibly. We discuss the "publish only" outcome in Why Not Everything Ethen Researches Needs to Become a Product.
On general intelligence
People sometimes ask whether Ethen is working toward artificial general intelligence. Our honest answer is that some of the long-horizon questions above — compositional skills, learning from experience, credit assignment, automated discovery — are part of what increasingly general intelligence would require. But no benchmark certifies general intelligence, and we assign no date, probability or guaranteed outcome to it. A credible contribution would be cumulative: reliable task execution, then recovery and transfer, then better learning from experience, then grounded models of action, each established with evidence before the next. If that work produces real advances, they will be reported as what they are.
Where the evidence stands
As of October 2026, Ethen Research Lab has 41 publications. One system card reports results on one pinned build. Thirteen publications synthesize existing evidence. The rest are proposals, protocols and benchmark designs. No publication is labeled as reporting new measured results more broadly. That is an accurate picture of a young research program that publishes its questions before its answers.
How to engage with the research
Researchers, practitioners and critics can help in practical ways. Reading a protocol and pointing out a flaw in its design before it runs is one of the most valuable contributions anyone can make, because flaws found early are cheap to fix. Suggesting a stronger baseline, a better metric or a dataset with clearer rights improves the eventual result. And asking for the evidence behind any Ethen claim — research or product — is always welcome; the evidence status on every publication is there to make that easy.
We also expect some of our questions to turn out to be the wrong ones. Part of publishing questions in public is being visibly wrong sometimes, and correcting course with an explanation rather than quietly.
Tradeoffs and limitations
Long horizons are uncertain. Many of these questions may not yield useful answers for years, or at all.
Breadth costs focus. Eight programs are a lot for a young lab. Not all are equally active, and priorities will change with evidence.
Questions are not progress. Publishing a question is the start, not an achievement.
No results claimed. Nothing here should be read as a finding.
FAQ
What does Ethen Research Lab research? How AI systems can learn from work that can be checked, across eight programs: evaluation and verification, trust and accountable work, adaptive intelligence, context and transfer, model intelligence, data and learning, enterprise and sovereign AI, and AI security.
What long-term questions is Ethen studying? Whether agents can predict the effects of their actions, combine skills on new tasks, learn which decisions led to success over long tasks, and help with research while keeping hypotheses separate from findings.
Is Ethen working on AGI? Some long-horizon questions relate to increasingly general intelligence, but Ethen sets no date or probability for it and makes no claim that any benchmark certifies it.
Has Ethen Research Lab published results? One system card reports results on one pinned build. Most publications are syntheses, proposals, protocols and benchmark designs.
Does all Ethen research become a product? No. Building, publishing only and stopping are all normal outcomes.
Related reading
- Why Ethen Research Lab Publishes Its Work in Public
- How to Explore Ethen Research Lab
- Why Not Everything Ethen Researches Needs to Become a Product
- Building a Research Lab as a Small Team in the Age of AI Agents
- Ethen Research Lab
References
- Gottweis, J., Weng, W.-H., Daryin, A., Tu, T., et al. (2025). Accelerating scientific discovery with Co-Scientist (originally "Towards an AI co-scientist"). arXiv:2502.18864. https://arxiv.org/abs/2502.18864
- Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., & Ha, D. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292. https://arxiv.org/abs/2408.06292
- Ethen Research Lab (2026). Verified Adaptive Intelligence: Learning From Work That Can Be Proven. Position paper. https://upcube.ai/resources/research/verified-adaptive-intelligence
- Ethen Research Lab (2026). How Should We Measure the Reliability of LLM Verifiers? Research protocol; not yet run. https://upcube.ai/resources/research/measuring-llm-verifier-reliability
- Ethen Research Lab (2026). The Capability Transfer Ledger. Methods paper; results gated. https://upcube.ai/resources/research/capability-transfer-ledger
- Ethen Research Lab (2026). Private AI Improvement Without Raw Data Export. Research proposal. https://upcube.ai/resources/research/private-ai-improvement