Skip to content

EthenEthenEthen

Ethen Code: What We're Building Next

Ethen Code is moving from assistive coding — explaining code, writing functions, fixing snippets — toward software tasks that end with evidence a person can check. The direction has seven parts: reproduce a problem before changing anything; plan a minimal, reviewable change before executing it; work in isolated environments with scoped permissions; treat tests, builds and review as evidence rather than as a finish line; put approvals in front of merges, deployments and other consequential steps; recover from failures in long-running work without repeating effects; and keep the same task model across Chat, the cloud workspace and Desktop. This article describes that direction. It is not a release schedule, and it makes no availability or date claims.

Ethen Code is moving from assistive coding — explaining code, writing functions, fixing snippets — toward software tasks that end with evidence a person can check. The direction has seven parts: reproduce a problem before changing anything; plan a minimal, reviewable change before executing it; work in isolated environments with scoped permissions; treat tests, builds and review as evidence rather than as a finish line; put approvals in front of merges, deployments and other consequential steps; recover from failures in long-running work without repeating effects; and keep the same task model across Chat, the cloud workspace and Desktop. This article describes that direction. It is not a release schedule, and it makes no availability or date claims.

Key takeaways

  • The unit of work is a task, not a prompt. Ethen Code is being built around tasks that state what "done" means and end with checkable evidence.
  • Reproduce first, then change. Showing the problem exists is the cheapest way to avoid fixing the wrong thing.
  • Plans before effects. Ethen Code is designed around plan-first execution with approval gates for consequential steps.
  • Passing tests are evidence, not proof. Tests show what they test; they do not prove security or business impact.
  • One system, three depths. Chat, the Platform workspace and Desktop are meant to share one task model, with local and cloud models as a per-task choice.

Where does Ethen Code stand today?

Ethen Code is designed as one coding system exposed in three places. In Chat, it is lightweight conversational help: explain this code, write a function, fix this bug. In Platform, it is a full cloud engineering workspace: repositories, branches, agents, terminal, multi-file editing, pull requests, builds, tests and environments. In Desktop, it is a local coding-agent environment: local repositories, local terminal, local models alongside Ethen's cloud models, Git and commands. The full description, with what is direction and what is implemented, is in Three Places to Use Ethen Code.

One implemented detail shows the kind of care the direction depends on. Coding run history is stored per signed-in account on both the web and desktop paths, cleared on sign-out, and isolated between accounts on the same device. That is a small thing, but trust in a coding agent starts with knowing whose work you are looking at. It is described in Keeping Code History Scoped to the Signed-In User.

What changes: from writing code to completing a task you can check

The next phase of Ethen Code is defined by a change in the unit of work. A prompt asks for code. A software task asks for an outcome — a bug fixed, a dependency updated, a test repaired — and ends with evidence that the outcome was reached.

That distinction matters because coding agents are now capable of long, multi-step work. Research tracking the length of software tasks that AI systems can complete has found that length growing quickly (Kwa et al., 2025). Longer tasks mean more places for a small early mistake to compound, and more temptation for an agent to report success at the end of a long run. The answer is not to make the agent sound more confident. It is to change what the agent is required to hand back.

Six boxes in two rows: Understand the request, Reproduce first, Plan a minimal change, Change in isolation, Check with evidence, Explain and hand off.
Figure 1. The direction for Ethen Code: a coding task ends with evidence a reviewer can check, not with a confident message.

1. Reproduce before you change

The first step of a good software task is showing that the problem exists. For a bug, that means a failing test or a reproducible sequence of steps. For a missing feature, it means a check that currently fails and should pass afterward.

Reproduction does three things. It confirms that the agent understood the request; a reproduction that does not match what the person described is caught before any code changes. It gives the change a target, so success can be measured rather than asserted. And it becomes part of the evidence the agent hands back: here is the failure, here is the change, here is the same check passing.

Not every task can be reproduced cleanly. Performance problems, flaky behavior and issues that depend on production data are hard to recreate. In those cases the direction is for Ethen Code to say so explicitly — "could not reproduce; proceeding on the description" — rather than to proceed as if a reproduction existed.

2. Plan a minimal, reviewable change

Ethen Code is designed around plan-first execution. Before editing files, the agent proposes what it intends to change and why. Plans serve people more than they serve models: a short plan is the fastest way for a reviewer to catch a misunderstanding, and a plan that touches twelve files for a one-line bug is a signal before any of those files change.

Minimal changes matter for the same reason. A small diff is easier to review, easier to revert and less likely to introduce unrelated regressions. The direction is for the agent to prefer the smallest change that makes the reproduction pass, and to separate refactoring from fixes when both are needed.

Where a step has consequences beyond the repository — merging into a protected branch, triggering a deployment, changing infrastructure configuration — the plan is where approval is requested. We describe why approvals sit in front of consequential actions, rather than everywhere, in Why Ethen Keeps Human Approval in the Loop.

3. Work in isolated environments with scoped permissions

A coding agent that runs commands needs somewhere safe to run them. The direction for the Platform workspace is that agents work in isolated environments: a copy of the repository and its dependencies, with the permissions the task needs and no more. Isolation protects shared systems from mistakes, keeps one task's side effects from leaking into another, and makes a task reproducible — the same environment can be restored to check or redo the work.

Scoped permissions are the other half. A task to fix a unit test does not need production credentials, and should not have them. The general principle, that agents should hold only the authority a task requires and that authority should be explicit, runs through Ethen's broader design and through the research proposal Mandates: Compiling Human Intent Into Bounded Agent Authority — a proposal that has not been implemented or tested as described.

4. Treat tests and review as evidence, not as a finish line

Tests passing is the most common way coding agents declare victory. It is also one of the most commonly misunderstood. A passing test suite shows that the code satisfies the tests. It does not show that the tests were adequate, that the change is secure, or that it does what the person actually wanted.

Ethen Code's direction is to hand back tests and review as evidence with a scope: which checks ran, which passed, which were added for this task, what the diff contains, and what was not checked. A reviewer can then decide how much more checking the change needs. The executable environments used by coding benchmarks make a similar point from the research side: grading by repository tests is powerful precisely because tests are executable, and it is limited precisely because tests can be incomplete (Jimenez et al., 2023; Pan et al., 2024).

This is one concrete case of a broader question we discuss in What "Done" Should Mean for an AI Agent: the agent finishing its steps is not the same as the task succeeding.

5. Approvals before consequential steps

Most coding work is reversible. A branch can be deleted; a commit can be reverted. Some steps are not, or are expensive to undo: a deployment that customers see, a migration that changes data, a message to a team, a change to access controls. The direction is for Ethen Code to distinguish those steps clearly and to put a human decision in front of them, bound to the exact change being approved.

Binding matters. An approval for "deploy this commit" should not quietly cover a different commit produced by a retry. Ethen already binds approvals to specific actions in its computer-use path, as described in Binding Computer-Use Approvals to Specific Actions. The same principle is the direction for consequential steps in code.

6. Recover from failures in long-running work

Longer coding tasks fail in the middle: a build server times out, a dependency download stalls, a test environment crashes, a deployment returns no confirmation. The direction is for Ethen Code tasks to run on the same durable foundation as other long-running Ethen work, so that a task can pause, resume from a known point, and avoid repeating an action whose outcome is uncertain.

Deployment is the clearest example. If a deploy request times out, the deploy may or may not have happened. Retrying blindly can deploy twice or roll forward something that should have been checked. The right first move is to look at the deployment system and find out. Ethen Research Lab's note Unknown Effects in Autonomous AI Systems sets out the general rule — reconcile before retrying — as a research synthesis. How we think about recovery as a product property is in Why Ethen Is Building for Recoverable AI Work.

7. Local and cloud, one task model

Developers move between their own machine and shared environments constantly. The direction is that a task in Ethen Code means the same thing whether it runs in the Platform workspace or on Desktop: the same plan, the same checks, the same evidence. Desktop adds local repositories, a local terminal and the option to use local models, with Ethen's cloud models available when a task needs them. The choice of where the model runs becomes a per-task decision rather than a choice of product.

We are explicit that moving in-progress work smoothly between local and cloud is direction, not a shipped capability. The groundwork — a shared run model stored per user on both paths — exists; synchronized workspace state does not yet.

8. Skills and conventions that survive model changes

Teams build up know-how about their codebases: how to run the tests, which modules are fragile, which conventions matter. Coding agents can use that know-how as instructions or reusable skills. The risk is that those instructions are tuned for one model and degrade silently when the model changes.

Ethen Research Lab has proposed separating what a skill promises from how a particular model is told to do it, in Skill IR: Toward Model-Independent Agent Capabilities, and has specified a protocol for measuring whether skills survive a frontier-model upgrade. Both are research — a proposal and a protocol not yet run — and neither describes current product behavior. The product direction they inform is simple: a team's coding conventions should be checked, not assumed, when the underlying model changes.

Table of Ethen Code's three surfaces — Chat, Platform workspace, Desktop — with what each is best for and its direction.
Figure 2. One Code system, three depths. The direction is the same task model across all three.

What we are not promising

Direction articles are easy to over-read, so here is what this one does not say.

  • No prompt-to-merged-pull-request promise. Moving from a request to a merged change crosses planning, checks and human review by design.
  • No autonomous deployment. Consequential steps are designed to require approval.
  • No dates or availability. This describes where Ethen Code is going, not when any part is available to any customer.
  • No benchmark scores. We have not published coding-agent benchmark results. Ethen Research Lab's VerifiedWork benchmark design includes code environments, but it has not been run.

How will we know it is working?

The measure of this direction is not how much code Ethen Code writes. It is whether tasks end in outcomes people accept without rework. The signals we care about are the ones that track that: changes accepted after review, changes reverted later, regressions introduced, time a reviewer spends understanding a change, and how often a task correctly stops and asks rather than guessing. Some of these take weeks to observe — a change can pass review and be reverted a week later — which is why we treat early success as provisional.

Tradeoffs

Reproduce-first and plan-first workflows are slower for trivial tasks. A one-line typo fix does not need a reproduction and a plan, and the product has to make the lightweight path lightweight. Isolated environments cost setup time and compute. Approvals add friction, and too many of them train people to approve without reading. The direction is to apply each control where its value is highest — consequential steps, multi-file changes, long-running tasks — and to keep conversational help in Chat fast.

Frequently asked questions

Is Ethen Code a coding agent or an editor? Both, at different depths. Chat offers conversational help, the Platform workspace offers a cloud engineering environment with agents, and Desktop offers a local coding-agent environment.

Can Ethen Code use local models? On Desktop, the design pairs local models with Ethen's cloud models so the choice can be made per task. Local model support differs by runtime.

Will Ethen Code deploy for me? Deployments are consequential steps. The direction is for them to require an explicit approval bound to the exact change.

Where should I start today? If your work fits in a message, start in Chat. If it needs branches and tests, use a workspace. Our guide Choosing an AI Coding Environment walks through the decision.

References

  1. Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. https://arxiv.org/abs/2503.14499
  2. Jimenez, C. E. et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv:2310.06770. https://arxiv.org/abs/2310.06770
  3. Pan, J. et al. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv:2412.21139. https://arxiv.org/abs/2412.21139
  4. Ethen Blog (2026). Three Places to Use Ethen Code. https://upcube.ai/blog/three-places-to-use-ethen-code
  5. Ethen Blog (2026). Keeping Code History Scoped to the Signed-In User. https://upcube.ai/blog/keeping-code-history-scoped-to-the-signed-in-user
  6. Ethen Blog (2026). Binding Computer-Use Approvals to Specific Actions. https://upcube.ai/blog/binding-computer-use-approvals-to-specific-actions
  7. Ethen Research Lab (2026). Skill IR: Toward Model-Independent Agent Capabilities. Research proposal; untested. https://upcube.ai/resources/research/skill-ir
  8. Ethen Research Lab (2026). Mandates: Compiling Human Intent Into Bounded Agent Authority. Research note; architecture proposal. https://upcube.ai/resources/research/agent-mandates