How AI Agents Are Changing the Way We Build Ethen
Using AI agents in software development has changed how Ethen is built, but not in the way people usually expect. The biggest change is not raw speed. It is where human attention goes. Agents now do much of the mechanical work — exploring unfamiliar code, implementing scoped changes, writing and running tests, and drafting reports of what they did. People spend more of their time defining tasks precisely, reviewing the evidence an agent produces, and deciding what is claimed, merged and deployed. Every agent-made change has to carry an evidence report saying what ran, what passed and what did not run, using the same status words as our release certificates. And no agent merges or deploys to production without explicit human authorization. This article describes that practice, what the research says about AI-assisted development, and what we have learned.
Using AI agents in software development has changed how Ethen is built, but not in the way people usually expect. The biggest change is not raw speed. It is where human attention goes. Agents now do much of the mechanical work — exploring unfamiliar code, implementing scoped changes, writing and running tests, and drafting reports of what they did. People spend more of their time defining tasks precisely, reviewing the evidence an agent produces, and deciding what is claimed, merged and deployed. Every agent-made change has to carry an evidence report saying what ran, what passed and what did not run, using the same status words as our release certificates. And no agent merges or deploys to production without explicit human authorization. This article describes that practice, what the research says about AI-assisted development, and what we have learned.
Key takeaways
- Attention moves, not just speed. People scope tasks, review evidence and decide what ships.
- Every change carries evidence. What ran, what passed, what did not run.
- Status words beat "done". PASS, PARTIAL, NOT_RUN, NOT_PERFORMED and OPEN.
- Claims stay scoped. Evidence attaches to a specific artifact and date.
- Humans authorize merges and deploys. Always explicitly.
- Productivity gains must be measured. Research shows feeling faster is not being faster.
What agents do in our engineering work
Coding agents have improved rapidly. Benchmarks such as SWE-bench, which tests whether models can resolve real issues from open-source repositories, have tracked that progress, and agents now handle many engineering tasks that once required a person throughout. At Ethen, agents routinely:
- Explore code. Read an unfamiliar part of the codebase, trace how something works, and summarize it.
- Implement scoped changes. Make a change described precisely, in an isolated copy of the code.
- Write and run tests. Add tests for the change and run the existing checks.
- Draft reports. Describe what was changed, what was checked, and what was not.
That covers a large share of day-to-day engineering. It does not cover deciding what to build, how the system should be structured, which claims are true, or what should reach users.
An agent-made change, start to finish
Figure 1 shows the workflow we use.
Scoped task. A person states the goal, the limits — what may be changed, what must not be touched — and what done means. Vague tasks produce vague changes.
Isolated work. The agent works on a separate copy of the code, never directly on production systems or shared environments.
Checks run. Type checks, builds, tests and policy checks run on the change.
Evidence report. The agent reports what it changed, which checks ran, which passed, which failed and which did not run — and why. This step is highlighted because it is the one that makes everything after it possible.
Human review. A person reviews both the change and its evidence. Reviewing evidence is often faster than reviewing every line, and it catches a different class of problem: a change that looks fine but was never actually tested in the way that matters.
Merge and deploy. Only with explicit human authorization. Agents do not merge or deploy on their own judgment.
Where agents carry the work, and where people stay
Figure 2 divides the responsibilities.
Agents do most of the exploring, implementing and testing. People check their conclusions, define scope, and decide what must be tested. For claims about what works, agents draft and people approve. Architecture and ownership decisions — which app owns which capability, how systems fit together — stay with people, with agents proposing options. Merging and deploying is never done by an agent alone.
Status words instead of "done"
The single most useful practice we have adopted is refusing to accept "done" as a report. An agent that says it has finished a task has made a claim. To review that claim, we need to know what was checked.
So agent reports use the same status vocabulary as Ethen's release certificates, shown in Figure 3.
PASS means the check ran and succeeded, on a named artifact and date. PARTIAL means some of the scope was verified and the rest is listed. NOT_RUN means a check exists but was not run this time. NOT_PERFORMED means the scope was deliberately out of bounds. OPEN means a known issue or decision remains unresolved.
We explain how to read these in published certificates in What a Release Certificate Actually Proves at Ethen. The same discipline applies inside engineering: a report full of PASS with no NOT_RUN entries is either a very complete test run or a report that is not telling the whole story, and reviewers learn to ask which.
Scoped claims, not general ones
Evidence attaches to a specific thing: a commit, a deployment, a date. A test that passed on one version of the code says nothing about the next version. A check that passed in a test environment says nothing about production.
Agents are prone to generalizing. "The tests pass" becomes "the feature works", which becomes "the feature is ready". Keeping claims scoped — this check, on this artifact, on this date — is a habit we enforce in review. When Ethen Designer moved into its own app, the published account was explicit about what recertification re-proved and what it explicitly did not do; see Moving Ethen Designer Into Its Own App. That is the standard we hold agent reports to.
Verify the state, not the summary
A practice that became important as agents took on more work: when picking up a task, check the actual state of the code and systems rather than trusting a summary of it. Summaries — including an agent's own notes from earlier work — go stale. Code changes, other work lands, environments drift.
Before an agent continues a task, it should confirm what is actually there: which files exist, which changes are already present, which checks currently pass. This catches a whole category of errors where work is redone, overwritten or built on assumptions that are no longer true.
A worked example
The following example is illustrative of the workflow rather than a record of a specific change.
An engineer asks an agent to add a retry limit to a background job, so that a job that keeps failing stops after a set number of attempts and is marked for review. The task states the goal, the files the agent may change, that existing behavior for successful jobs must not change, and that done means: the limit is enforced, tests cover the new behavior, and all existing checks pass.
The agent explores the job code, finds where attempts are counted, implements the limit and adds tests. It runs the type checks, the build and the test suite. Its report lists: type checks PASS; build PASS; new tests PASS; existing job tests PASS; an integration test that requires a live queue NOT_RUN, because no queue is available in the isolated environment; and OPEN, a question about whether jobs that hit the limit should notify anyone.
The engineer reviews the change and the report. The NOT_RUN line matters: the behavior has not been tested against a real queue. They run the integration test in a staging environment, which passes, and decide that notification is a separate task. They then authorize the merge. The deployment happens later, as a separate, explicit decision.
Without the status words, the agent's report would have said "done", and the untested integration path would have been invisible.
Access and security for agents
Agents that write code need access to code, and that access needs the same care as any other credential.
Least access. Agents get access to what a task needs: the relevant code, the ability to run checks, and nothing more. They do not hold production credentials.
No secrets in prompts or reports. Credentials are never pasted into agent instructions, and agent reports must not echo them.
Separate environments. Agents run checks in isolated environments. Anything that touches shared systems goes through the same review and authorization as human changes.
Untrusted inputs stay untrusted. Code comments, documentation and issue text can contain instructions. Agents treat them as information about the task, not as commands, and changes that seem to follow such instructions get extra scrutiny in review.
What the research says about productivity
We are careful about productivity claims, because the evidence is mixed.
A 2023 controlled experiment by researchers at GitHub and collaborators found that developers using an AI pair programmer completed a specific, well-defined task — implementing a small web server — about 56% faster than a control group. That is a large effect on a bounded task.
A 2025 randomized controlled trial by METR found something very different. Experienced open-source developers working on their own large, familiar projects, using early-2025 AI tools, took about 19% longer with AI assistance than without. Strikingly, they had predicted beforehand that AI would make them faster, and still believed afterward that it had.
Both results can be true. AI assistance can help a lot on well-defined tasks in unfamiliar territory and help little — or slow people down — on complex work in code they already know well, where checking and correcting AI output costs more than it saves. The practical lesson is to measure on your own work rather than rely on the feeling of speed. We have not published productivity measurements for Ethen's engineering, and we are not claiming any here.
What we have learned
Precise tasks matter more than ever. Agent output quality depends heavily on how clearly the task, limits and done-criteria are stated.
Review shifts from lines to evidence. Reviewing what was checked often finds more than reviewing every line.
Agents are confident. An agent's report reads as confident whether or not it is right. Status words force the uncertainty into view.
Isolation prevents accidents. Working on separate copies, never on shared or production systems, means mistakes stay contained.
Authorization must be explicit. "Looks good" in a conversation is not authorization to deploy. Deploying needs a deliberate decision, every time.
Ownership questions do not go away. Agents can implement anything they are asked to; deciding which app owns which capability, and keeping that consistent, is still a human design job.
How this connects to Ethen Code
The practices we use internally shape the product we build for others. Ethen Code is designed around the same ideas: scoped work, checks that run, evidence of what was and was not verified, and humans in control of what ships. We describe the direction in Ethen Code: What We're Building Next and the surfaces where Ethen Code runs in Three Places to Use Ethen Code. For choosing between AI coding setups more generally, see Choosing an AI Coding Environment.
Tradeoffs and limitations
Evidence reports take time. Writing and reviewing them adds overhead to every change.
Agents can produce plausible but wrong evidence. A report can claim a check passed when it did not run as described. Reviewers spot-check the evidence itself.
Not every task suits an agent. Deep architectural work and subtle debugging in familiar code may be faster done directly.
No productivity claims. We describe practice, not measured gains.
FAQ
How do engineering teams use AI coding agents safely? By scoping tasks precisely, having agents work in isolated copies, requiring evidence reports of what ran and did not, reviewing that evidence, and keeping merges and deploys under explicit human authorization.
Do AI coding agents make developers faster? Sometimes. A controlled experiment found large gains on a well-defined task; a 2025 trial found experienced developers slower on their own complex projects. Measure on your own work.
How should AI-written code be reviewed? Review the change and its evidence: which checks ran, which passed, which did not run, and whether claims are scoped to the right artifact.
Does Ethen let AI agents deploy to production? No. Merges and deploys require explicit human authorization.
What are release certificate status words? PASS, PARTIAL, NOT_RUN, NOT_PERFORMED and OPEN — each describing exactly what a check did or did not establish.
Related reading
- What a Release Certificate Actually Proves at Ethen
- Moving Ethen Designer Into Its Own App
- Ethen Code: What We're Building Next
- Choosing an AI Coding Environment
- Building a Research Lab as a Small Team in the Age of AI Agents
References
- Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089. https://arxiv.org/abs/2507.09089
- Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590. https://arxiv.org/abs/2302.06590
- Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024. arXiv:2310.06770. https://arxiv.org/abs/2310.06770
- Ethen Blog. What a Release Certificate Actually Proves at Ethen. https://upcube.ai/blog/what-a-release-certificate-actually-proves-at-ethen