Skip to content

EthenEthenEthen

Building a Research Lab as a Small Team in the Age of AI Agents

A small AI research lab can now do work that once needed a much larger team, because AI agents can search and summarize literature, draft and restructure documents, check consistency and citations, produce figures and scaffold experiments. That is how Ethen Research Lab operates. But agent output is not evidence, and a lab that produces more than it can check produces volume, not knowledge. So the operating model rests on three commitments. People own the questions, the judgment about evidence, the methods fixed before results, and the decision to publish. Agents do bounded labor inside those decisions. And the safeguards — a registry of checked sources, automated checks, review by someone who did not produce the work, protocols published before experiments, and evidence labels on every publication — have to scale as fast as the output does. This article explains that model, where agents help, where they mislead, and what we would advise another small lab.

A small AI research lab can now do work that once needed a much larger team, because AI agents can search and summarize literature, draft and restructure documents, check consistency and citations, produce figures and scaffold experiments. That is how Ethen Research Lab operates. But agent output is not evidence, and a lab that produces more than it can check produces volume, not knowledge. So the operating model rests on three commitments. People own the questions, the judgment about evidence, the methods fixed before results, and the decision to publish. Agents do bounded labor inside those decisions. And the safeguards — a registry of checked sources, automated checks, review by someone who did not produce the work, protocols published before experiments, and evidence labels on every publication — have to scale as fast as the output does. This article explains that model, where agents help, where they mislead, and what we would advise another small lab.

Key takeaways

  • Agents expand labor, not judgment. People decide what to ask, what counts as evidence and what to publish.
  • Output is not evidence. More papers are only better if checking keeps pace.
  • Independence is non-negotiable. Review by someone, or something, that did not produce the work.
  • Fix methods before results. Protocols published first keep fast-moving work honest.
  • Label everything. Every publication states what kind of evidence it holds.
  • Measure, do not assume, productivity. Research on AI-assisted work shows gains are not automatic.

What changes for a small lab

Research labs have traditionally scaled with people. Literature reviews, drafting, figure preparation, experiment setup and consistency checking all take time, and a small team runs out of it quickly. AI agents change that arithmetic. A small team can now survey a literature, draft a series of papers, keep terminology consistent across them, and produce clean figures in a fraction of the time it once took.

Ethen Research Lab's 40-paper release, described in What We Learned Publishing 40 Research Papers at Once, would not have been feasible for a small team without that help. It also taught us that the bottleneck moves. When producing text becomes cheap, the scarce resources become judgment, checking and honesty about what has and has not been shown.

Who does what

Figure 1 sets out the division of labor.

Three columns: people own (highlighted) questions, judgment, methods and sign-off; agents help with literature, drafting, checks, figures and scaffolding; neither alone should decide novelty, interpret surprising results, or decide what not to publish.
Figure 1. Agents expand the labor; people keep the judgment.

People own the choice of questions, judgment about what counts as evidence, methods fixed before results, and sign-off to publish. These are where research succeeds or fails, and they cannot be delegated.

Agents help with searching and summarizing literature, drafting and restructuring, checking consistency and citations, producing figures, and scaffolding code for experiments. These are labor-intensive, and agents do them well when the work is bounded and checked.

Neither alone should claim novelty, interpret surprising results or decide what not to publish. These need the combination: agents can surface relevant prior work and alternative explanations; people decide.

Where agents help, and where they mislead

Every place an agent helps is also a place it can mislead. Figure 2 sets them side by side.

Table of five research tasks with where agents help and where they can mislead, with citations highlighted: plausible citations that do not support the claim.
Figure 2. Every row's right-hand column is a reason for a check.

Literature. Agents find and summarize many sources quickly. They can also misattribute a claim or number to the wrong source.

Drafting. Agents produce clear structure and prose fast. Fluent text can easily say more than the evidence supports.

Citations. This row is highlighted because it is the most dangerous. A study of generative search engines by Nelson Liu, Tianyi Zhang and Percy Liang found that only about half of generated sentences were fully supported by the citations attached to them. In research writing, a plausible citation that does not actually support the claim is worse than no citation, because it looks like evidence.

Experiments. Agents can scaffold code and runs. A silent bug can produce clean-looking results that are simply wrong.

Figures. Agents produce consistent diagrams quickly. A chart drawn without data can imply results that do not exist.

The automated-research literature shows both sides. The "AI Scientist" project demonstrated language models carrying out machine learning research from idea to paper, and also relied on automated review rather than human experts — an illustration of how quickly fluent output can outrun independent checking.

Safeguards that scale with output

If output grows and checking does not, quality falls. Figure 3 shows the safeguards we rely on.

Five safeguards that scale with output: source registry, automated checks, independent review (highlighted), methods before results, and public evidence labels.
Figure 3. If output grows faster than checking, a lab produces volume, not knowledge.

A source registry. Every external claim traces to a source whose content was checked at the time of writing. Numbers are attributed only when the source actually states them. Sources that cannot be confirmed are not cited.

Automated checks. Rules that can be stated precisely — required sections, working links, numbered references, evidence labels, figure order, banned overclaiming words — run on every document. We describe these in our 40-paper lessons.

Independent review. Work is reviewed by someone, or a separate review session, that did not produce it. Small teams often combine roles; that is acceptable only if reviewer independence is preserved. The same person or process that wrote a paper should not be the only one to approve it.

Methods before results. Protocols are fixed and published before experiments run, so results cannot quietly reshape the method. This follows the preregistration practice advocated by researchers concerned with selective reporting.

Public evidence labels. Every publication states what kind of evidence it holds, so a reader — or an AI system summarizing the page — gets the right qualifier.

A worked example: one publication, start to finish

The following example is illustrative of the operating model rather than a record of a specific paper.

A researcher decides the Lab should ask whether a particular recovery strategy helps agents finish long tasks. That decision — the question, and why it matters — is theirs.

An agent searches the literature, summarizes a few dozen relevant papers, and proposes how the question relates to existing work. The researcher reads the most relevant sources directly, adds two the agent missed, and removes one that was only superficially related.

The researcher drafts the protocol's core: the hypothesis, the comparison, the measures and what result would count as failure. An agent helps turn that into a full document, checks terminology against the Lab's other publications, and produces figures for the experimental design.

Automated checks run: every reference in the registry, every link working, the evidence label present, no overclaiming language. One citation fails — the number attributed to it does not appear in the source — and is corrected.

A separate review session, with a checklist, reads the protocol as a skeptic would. An agent is also asked to list every claim that is not supported by a cited source. The reviewer finds one claim stated too strongly and softens it.

The researcher signs off. The protocol is published, labeled as a protocol not yet run. Only later, when the experiment has been run as specified, will results be published — with any deviations from the protocol stated.

Choosing questions a small lab can answer

Agents expand what a small team can produce, but they do not expand everything. Compute, access to data with clear rights, domain expertise and time for careful experiments remain limited. A small lab does best with questions where those limits bite least: questions about methods, measurement and evaluation; questions that can be studied in software environments that are cheap to reset; and questions where a careful negative result is valuable. Questions that need enormous compute or proprietary data at scale are usually better left to others, or approached through collaboration.

What we got wrong early

Two mistakes are worth sharing. First, we underestimated how much checking agent-assisted drafting needs. Early drafts read well, which made them easy to trust, and the source registry and automated checks repeatedly caught fluent sentences whose cited sources did not say what the sentence claimed. Second, we initially treated review as a final step. It works much better as a continuous one, with checks running on every draft rather than only at the end.

Productivity is not automatic

It is tempting to assume that AI agents simply make researchers faster. The evidence is more mixed. A 2025 randomized controlled trial by METR, studying experienced open-source developers working on their own projects with early-2025 AI tools, found that the developers took about 19% longer with AI assistance than without — even though they predicted beforehand, and believed afterward, that AI had sped them up.

That study concerned software development, not research, and tools have changed since. But its lesson applies: the feeling of speed is not a measurement of it. A small lab using agents should check whether they actually help on its tasks, and where they do not, stop using them for those tasks. The same caution applies to our own claims about how AI agents change our work, discussed in How AI Agents Are Changing the Way We Build Ethen.

Independence when the team is small

Reviewer independence is the hardest safeguard for a small team. When three people do everything, each piece of work has few independent readers. Several practices help.

Separate production and review sessions. Even when the same person is involved, reviewing in a separate session, with fresh eyes and a checklist, catches more than reviewing while writing.

Use agents as adversarial reviewers. An agent asked to find unsupported claims, missing limitations or misattributed numbers is a useful first pass — not a substitute for a human reviewer, but a way to make human review more targeted.

Bring in outside experts. A domain expert or independent scientific reviewer adds more credibility to a small lab than a ceremonial advisory board. Collaboration with universities on methods, benchmarks and evaluation adds independence too.

Publish protocols for criticism. A protocol published before an experiment invites anyone to point out flaws while they can still be fixed.

How collaboration fits

A small lab cannot be expert in everything, and agents do not supply expertise they were never given. Collaboration fills that gap. Working with university researchers on open methods, benchmarks and evaluation brings independent judgment and domain depth. Agreements need to be clear from the start about publication rights, data rights, intellectual property and reproducibility, so that collaboration strengthens independence rather than complicating it. Publishing protocols openly also makes informal collaboration possible: anyone who spots a flaw in a published method is, in effect, a reviewer.

What we would advise another small lab

Decide what you will not claim before you start. Write down what evidence each kind of claim needs.

Build the source registry first. It is much harder to verify citations after the fact.

Automate the boring checks early. They pay for themselves within a handful of documents.

Keep one person accountable for each publication. Agents help; a named person signs off.

Publish questions and methods, not just results. It disciplines the work and invites help.

Measure whether agents help. On your tasks, with your team, rather than assuming.

Tradeoffs and limitations

Speed creates temptation. When drafting is cheap, it is tempting to publish more than can be checked. The safeguards exist to resist that.

Agents introduce new failure modes. Misattributed citations and fluent overclaiming are more common with agent drafting, not less.

Independence is hard to guarantee in a small team. Outside review helps but is not always available.

This describes an operating model. It is not a claim about productivity gains, which we have not measured.

FAQ

Can a small team run a serious AI research lab? Yes, if it uses agents for bounded labor, keeps judgment and sign-off with people, and scales its checking as fast as its output.

How should researchers use AI agents? For literature search, drafting, consistency checks, figures and experiment scaffolding — with every claim traced to a checked source and every result independently reviewed.

How do you keep AI-assisted research honest? A source registry, automated checks, independent review, methods fixed before results, and evidence labels on every publication.

Do AI agents make researchers faster? Not automatically. A 2025 trial found experienced developers slower with early-2025 AI tools despite believing they were faster. Measure on your own tasks.

Does Ethen Research Lab use AI agents? Yes, for bounded research labor, with people owning questions, judgment and publication decisions.

References

  1. Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089. https://arxiv.org/abs/2507.09089
  2. Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating Verifiability in Generative Search Engines. Findings of EMNLP 2023. arXiv:2304.09848. https://arxiv.org/abs/2304.09848
  3. Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., & Ha, D. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292. https://arxiv.org/abs/2408.06292
  4. Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. https://doi.org/10.1073/pnas.1708274114