Skip to content

EthenEthenEthen

What It Means to Build an AI Company With AI

An AI-native company is one that uses AI agents throughout its own work — engineering, research, writing, analysis and operations — rather than only selling AI to others. Ethen is built that way. But using agents everywhere is the easy part. The hard part is keeping the company honest while it does: making sure people stay accountable for every decision, agents work within bounded authority, the context they rely on is written down where both people and agents can read it, claims carry evidence rather than confidence, output and spend are measured rather than assumed, and the judgment that decides what is right is hired for, practiced and protected. Those six principles are the same ones we design into Ethen's products, applied to ourselves. This article explains them, where AI helps and misleads inside a company, and the risks we watch for.

An AI-native company is one that uses AI agents throughout its own work — engineering, research, writing, analysis and operations — rather than only selling AI to others. Ethen is built that way. But using agents everywhere is the easy part. The hard part is keeping the company honest while it does: making sure people stay accountable for every decision, agents work within bounded authority, the context they rely on is written down where both people and agents can read it, claims carry evidence rather than confidence, output and spend are measured rather than assumed, and the judgment that decides what is right is hired for, practiced and protected. Those six principles are the same ones we design into Ethen's products, applied to ourselves. This article explains them, where AI helps and misleads inside a company, and the risks we watch for.

Key takeaways

  • Accountability stays with people. Every decision has a named owner.
  • Agents get bounded authority. The same rule we build into products.
  • Write it down. Written context is how agents and people share understanding.
  • Evidence over confidence. Claims carry status words and scope.
  • Measure, do not assume. Feeling faster is not being faster.
  • Protect judgment. It is the scarce resource when output is cheap.

Why "AI-native" needs a definition

"AI-native" is used loosely. Sometimes it means a company whose product is AI. Sometimes it means a company that uses AI tools heavily. Sometimes it is just a label. We use it in a specific sense: a company that relies on AI agents to do a substantial share of its work, and that has therefore had to work out how to keep that work trustworthy.

That second part is what matters. Any company can give its staff AI tools. Building a company whose engineering, research and communication rely on agents — and whose claims to customers still hold up — requires deliberate principles. Figure 1 lists ours.

Six stacked principles: people stay accountable; agents get bounded authority; written context is the substrate; claims carry evidence (highlighted); measure, do not assume; protect judgment.
Figure 1. The principles we apply to our products, applied to ourselves.

Principle 1: people stay accountable

Every decision has a named person who owns it. Agents draft, implement, analyze and recommend. People decide what to build, what to claim, what to publish and what to ship.

This sounds obvious and is easy to erode. When agents produce most of the work, it becomes tempting to treat their output as the decision — to merge the change because the agent said it was done, publish the paper because it reads well, send the message because it was drafted. Naming an owner for every decision is the countermeasure. If no one can say who decided something, no one did.

Principle 2: agents get bounded authority

Agents inside Ethen work under the same principle we build into the product: authority bounded to the task. An agent changing code works on an isolated copy and cannot deploy. An agent drafting a document cannot publish it. An agent analyzing data sees the data the task needs.

We describe the product version of this principle in What We Mean by Responsible Autonomy. Applying it internally is both a safety measure and a test of our own thinking: if bounded authority is too cumbersome for us, it will be too cumbersome for customers.

Principle 3: written context is the substrate

Agents work from what they can read. Decisions made in conversation and never written down are invisible to them, and increasingly invisible to colleagues too. So the company runs on written context: plans, decisions with their reasons, product boundaries, evidence of what was checked, and lists of what remains open.

Writing things down has always been good practice. With agents, it becomes infrastructure. A clear written decision lets an agent work consistently with it; an unwritten one guarantees drift. The discipline also helps people: new team members, and existing ones returning to an old problem, find the reasoning instead of reconstructing it.

Principle 4: claims carry evidence

Inside Ethen, "done" is not a report. Claims carry status words — PASS, PARTIAL, NOT_RUN, NOT_PERFORMED, OPEN — scoped to a specific artifact and date, as our release certificates do; see What a Release Certificate Actually Proves at Ethen. Research publications carry evidence statuses. Engineering changes carry evidence reports, as described in How AI Agents Are Changing the Way We Build Ethen.

This matters more with agents than without them, because agent output is uniformly confident. A system that reports every result in the same assured tone gives reviewers nothing to work with. Status words force uncertainty into view.

Principle 5: measure, do not assume

AI assistance can make work dramatically faster, or slower, depending on the task. A large field experiment with management consultants, published as a working paper in 2023 by Fabrizio Dell'Acqua and colleagues, found that for tasks within AI's capabilities, consultants with AI completed more tasks, faster and at higher quality. For a task designed to fall just outside those capabilities, consultants with AI were markedly less likely to reach the correct answer than those without. The authors called this a "jagged technological frontier": tasks that look similar can fall on opposite sides of it.

A 2025 randomized trial with experienced software developers found a similar warning in a different setting: they were slower with AI assistance on their own complex projects, while believing they were faster.

So we try to measure whether AI helps on our specific tasks rather than assuming it does. Figure 2 summarizes how we treat different kinds of work.

Table of five kinds of work with how AI tends to perform and how Ethen responds, with judgment calls that look routine highlighted: AI can mislead confidently, so people keep the decision.
Figure 2. The frontier is jagged: tasks that look alike can fall on opposite sides of it.

Drafting, summarizing and restructuring are firmly inside the frontier; we use agents freely and review. Routine code changes with tests are too; we review the evidence. Judgment calls that look routine are the danger zone — AI can be confidently wrong — so people keep them and use AI as a second opinion. Novel analysis in familiar territory is where AI helps less than it feels like it does, so we measure. And decisions with external consequences require explicit human authorization.

Principle 6: protect judgment

When output becomes cheap, judgment becomes the scarce resource: knowing which question matters, which result is surprising, which claim is too strong, which design will age well. A company built with AI should hire for judgment and taste, give people time to exercise them, and keep review independent of production.

Automation research has a long-standing warning here. Lisanne Bainbridge's 1983 paper on the "ironies of automation" observed that the more a system automates, the harder the human's remaining role becomes — and the less practice people get at the skills they need when automation fails. A company that lets agents do everything risks staff who can no longer tell when the agents are wrong. So people keep doing some of the work themselves, deliberately.

A worked example: one week of work

The following example is illustrative of how the principles combine in ordinary work.

On Monday, a product lead writes a short decision: a new setting will let organizations turn off memory for a project. The decision names its owner, the reason and what is out of scope. That document is the context everything else draws on.

On Tuesday, an engineer gives an agent a scoped task based on the decision. The agent implements the setting in an isolated copy and reports: type checks PASS, unit tests PASS, an end-to-end test that needs a signed-in session NOT_RUN. The engineer runs the missing test in staging and reviews the change.

On Wednesday, a writer asks an agent to draft the documentation from the decision and the change. The draft describes the setting as removing all stored memory, which overstates it — the setting stops new memory but a separate action deletes existing memory. The writer corrects it, because they read the decision, not just the draft.

On Thursday, the lead authorizes the release. On Friday, a support analyst uses an agent to summarize early questions from customers and notices that several are confused about exactly the distinction the writer corrected. The documentation is updated with a clearer example.

At every step, an agent did most of the labor and a person made the call. The one error that could have reached customers — the overstated documentation — was caught because someone read the written decision rather than trusting a fluent draft.

Hiring and growing people in an AI-native company

If agents do much of the production, what should a company look for in people? We think three things matter more than before.

Judgment. The ability to tell a strong result from a weak one, a claim from evidence, and a good question from a busy one.

Clear writing. Because written context is the substrate, people who can state goals, limits and decisions precisely make everyone — human and agent — more effective.

Skepticism toward fluency. The habit of asking "how do we know?" of polished output, including one's own.

Risks we watch for

Figure 3 lists the risks we track.

Three columns of risks: quality (highlighted) — over-trust, errors outside the frontier, volume without review; people — skill atrophy, review fatigue, sameness; systems — over-broad access, secrets in prompts, untrusted instructions.
Figure 3. Each risk has a matching principle; none is solved once.

Quality risks. Over-trust in fluent output, errors on tasks outside the frontier, and volume that outruns review. Research on trust in automation describes the goal as calibrated reliance — trusting automation as much as its demonstrated reliability justifies, no more and no less.

People risks. Skill atrophy, fatigue from reviewing large volumes of agent output, and sameness — when everyone uses the same tools, their work can converge on the same ideas.

System risks. Agents with broader access than they need, credentials pasted into prompts, and untrusted text in inputs — documents, issues, web pages — that tries to instruct an agent.

Each risk maps to a principle: accountability and evidence against over-trust; measurement against frontier errors; protected judgment against atrophy and sameness; bounded authority against system risks. None is solved once; all need attention as tools and work change.

What changes, and what does not

Some things change dramatically when a company builds with AI. The cost of producing a first draft of almost anything — code, analysis, documentation, research summaries — falls sharply. Small teams can attempt work that once needed large ones; we describe what that means for research in Building a Research Lab as a Small Team in the Age of AI Agents.

Other things do not change at all. Customers still need claims they can rely on. Products still need someone to decide what they are for. Mistakes still have consequences. And the people accountable for a company's work are still people.

Building what we use

Building Ethen with AI agents has one more benefit: we are our own first users. The problems we hit — vague tasks producing vague results, agents claiming "done" without evidence, approvals that are too broad, context that was never written down — are the problems Ethen's products are designed to solve. Approvals bound to specific actions, evidence-gated completion, honest unknowns and delegated jobs with checked outcomes all reflect lessons we learned doing our own work. We describe the product versions in Why Ethen Founder Is About Outcomes, Not More Chat and Asking AI vs Delegating Work to AI.

Tradeoffs and limitations

Principles add overhead. Evidence reports, written decisions and independent review take time that pure speed would skip.

Measurement is hard. Knowing whether AI helps on a task requires comparison, which is often impractical. We measure where we can and stay skeptical where we cannot.

We are still learning. Tools change quickly, and practices that work today may need revising.

No productivity claims. This article describes principles, not measured gains.

FAQ

What does an AI-native company look like? One that relies on AI agents for a substantial share of its work while keeping people accountable, agents bounded, context written, claims evidenced, results measured and judgment protected.

How do you run a company with AI agents? Name an owner for every decision, scope agent authority to each task, write down context agents rely on, require evidence for claims, measure whether AI helps, and keep review independent.

What are the risks of relying on AI inside a company? Over-trust in fluent output, errors on tasks AI handles poorly, skill atrophy, review fatigue, sameness of thinking, over-broad agent access and manipulation through untrusted inputs.

What is the jagged technological frontier? A description, from a 2023 field experiment, of AI's uneven capabilities: tasks that seem similarly hard can fall inside or outside what AI does well.

Does Ethen use its own products? Yes. Many of Ethen's product principles come from lessons learned using AI agents in our own work.

References

  1. Dell'Acqua, F., McFowland III, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School Working Paper 24-013. https://papers.ssrn.com/abstract=4573321
  2. Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089. https://arxiv.org/abs/2507.09089
  3. Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775–779. https://doi.org/10.1016/0005-1098(83)90046-890046-8)
  4. Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80. https://doi.org/10.1518/hfes.46.1.50_30392