Skip to content

EthenEthenEthen

Why We’re Choosing Depth Over Feature Count

When we weigh product depth vs feature count, we choose depth. Feature count measures inventory: how many buttons, modes and models a product can list. Depth measures whether work actually finishes: whether a workflow covers the real task, works reliably on supported inputs, recovers when a step fails, returns a result you can check, keeps its context between sessions, and lets you control what it touches. AI products are unusually prone to feature sprawl, because every new model capability can become a new button in an afternoon. Ethen deliberately does fewer things per surface and tries to do each one completely. We give each kind of work one home, keep Ethen Chat focused, qualify creative workflows before exposing them, and ask a short set of public questions before adding anything. This article explains why, what we mean by depth, and what the choice costs.

When we weigh product depth vs feature count, we choose depth. Feature count measures inventory: how many buttons, modes and models a product can list. Depth measures whether work actually finishes: whether a workflow covers the real task, works reliably on supported inputs, recovers when a step fails, returns a result you can check, keeps its context between sessions, and lets you control what it touches. AI products are unusually prone to feature sprawl, because every new model capability can become a new button in an afternoon. Ethen deliberately does fewer things per surface and tries to do each one completely. We give each kind of work one home, keep Ethen Chat focused, qualify creative workflows before exposing them, and ask a short set of public questions before adding anything. This article explains why, what we mean by depth, and what the choice costs.

Key takeaways

  • Feature count is inventory, not value. A long list says what a product can show, not what it can finish.
  • AI makes sprawl cheap. New model capabilities are easy to wrap in a button and hard to make dependable.
  • Depth is checkable. Coverage, reliability, recovery, evidence, continuity and control can each be tested.
  • Specialized surfaces protect depth. Each kind of work has one home, so improvements accumulate in one place.
  • Saying "not now" is normal. Most proposals are either in the wrong place or not yet dependable.
  • Breadth still has a role. Reference surfaces such as a model library should be broad; workflows should be deep.

Why do AI products get cluttered?

AI products get cluttered because adding a capability is cheap and making it dependable is expensive. When a new model can generate audio, edit images or browse the web, a product team can expose that capability behind a new button almost immediately. The demo works. The release notes grow. What usually lags is everything around the capability: input validation, failure handling, cost controls, history, review, and a clear answer to "where does this live?"

Three forces make this worse in AI than in most software.

Model releases arrive constantly. Each one invites a new mode, toggle or model picker entry. A product that exposes every release eventually becomes a catalog rather than a tool.

Demos reward breadth. A one-minute demo can show twenty capabilities. It cannot show what happens on the fortieth run, when an input is malformed, or when a long job is interrupted halfway through.

Generation looks finished. A fluent paragraph or a striking image looks complete even when the surrounding work — checking, revising, delivering, recording what was done — has not happened. That makes it easy to mistake a capability for a workflow.

We have argued elsewhere that better models make the surrounding product more important, not less; see Why Better AI Models Don't Eliminate the Need for Better Products. Feature sprawl is the opposite bet: that the model is the product and the interface only needs to expose it.

Feature count versus depth

Feature count and depth answer different questions. Feature count answers "what can this product show me?" Depth answers "will this product finish my work, and can I tell whether it did?" Figure 1 summarizes the contrast.

Two columns. Feature count measures listed buttons, modes and models, demo appeal and inventory at purchase. Depth (highlighted) measures whether work finishes end to end, what happens on failure, whether the result can be checked, and whether the work persists.
Figure 1. Feature count describes inventory; depth describes finished work.

There is good evidence that the two diverge in practice. In a series of studies published in 2005, Debora Viana Thompson, Rebecca Hamilton and Roland Rust found that people choosing a product before use gave more weight to capability, so they tended to pick feature-rich options, while after using it they gave more weight to usability and became less satisfied with the complex option. The authors called this "feature fatigue." A product optimized to win the comparison table can lose the user.

Choice has costs too. A widely cited 2000 experiment by Sheena Iyengar and Mark Lepper found that shoppers offered a smaller assortment were more likely to make a purchase than shoppers offered a much larger one. Later research shows the effect varies by context, so we treat it as a caution rather than a law. But anyone who has scrolled through a model picker with dozens of near-identical entries has felt it.

What we mean by depth

Depth, in Ethen's sense, is a set of properties a workflow either has or does not have. We use six, shown in Figure 2.

Six stacked properties of a deep workflow: coverage, reliability, recovery, evidence (highlighted), continuity and control, each with a one-line definition.
Figure 2. Our working definition of depth. Each property can be checked, unlike a feature list.

Coverage. The workflow handles the real task, including its awkward parts. A video workflow that accepts a prompt but not a reference image covers the demo, not the job.

Reliability. It behaves the same way on the tenth run as the first, on the inputs it says it supports. Inputs it cannot handle are rejected up front, before anything is spent.

Recovery. When a step fails or its outcome is unclear, the work can resume or be reconciled instead of silently starting over. Our published engineering posts describe treating an unknown outcome as its own state rather than a reason to retry; see When an Agent Action's Outcome Is Unknown.

Evidence. The result arrives with a record of what was done and what was checked, so a person can review it rather than trust it.

Continuity. Context, artifacts and decisions persist across sessions, so the work is still there tomorrow.

Control. People can see and bound what the system may touch and spend.

Each property is testable. That is the main reason we prefer depth as a goal: you can tell whether you have it. "More features" can always be satisfied by adding another one.

Why specialized surfaces protect depth

Depth needs a home. If the same capability exists in three places, improvements land in one of them and the others drift. That is why Ethen is organized as a family of specialized apps — Chat, Code, Studio, Research, Designer, Founder, the Platform products and Desktop — with one canonical owner for each capability. We explain the company rationale in Why Ethen Is a Family of Specialized AI Apps, Not One App.

The clearest example is Ethen Chat. Chat is designed to stay deliberately limited: a fast conversational surface with lightweight entry points into deeper work, and hand-offs to the app that owns it. Research that accumulates sources belongs in Research. Creative projects with history and assets belong in Studio. Repository work belongs in Code. Our post Keeping Ethen Chat Focused describes those boundaries, and it is also candid that the current Chat model picker lists more entries than the curated design calls for. Holding a boundary is ongoing work, not a one-time decision.

The same principle shaped Ethen Studio. Rather than exposing every creative model behind a generate button, Studio's published media pipeline runs four qualified workflows — text-to-image, image editing, text-to-video and image-to-video — each on a verified route, with inputs validated, cost quoted and reserved before work runs, and uncertain outcomes reconciled rather than blindly retried. A model appearing in a catalog does not make it runnable. We describe where Studio is going in Why Ethen Studio Is Becoming More Workflow-Oriented.

How we say no

Most proposals we decline are not bad ideas. They are either in the wrong place or not yet dependable. Figure 3 shows the questions we ask. They are principles we are willing to publish, not a scoring formula.

Six questions in sequence: what work does it finish; which surface owns it; can we make it dependable; can we show evidence (highlighted); what does it cost everyone else; can we keep maintaining it. A note says 'not now' is the common answer.
Figure 3. The questions are public principles, not a scoring formula.

What work does it finish? We ask for a task and a person, not a model capability. "Generate audio" is a capability. "Produce a narrated draft of a product walkthrough that a marketer can revise" is work.

Which surface owns it? If the honest answer is "a little bit of everywhere," the capability is not ready. One surface should own it; others can link or hand off.

Can we make it dependable? Supported inputs, failure handling and recovery come before exposure. A capability that works only on the happy path is a demo.

Can we show evidence? If we cannot tell a person what was done and what was checked, we are asking them to trust fluency. We explain why that matters in What Ethen Is Doing to Make AI Outputs Easier to Verify.

What does it cost everyone else? Every option is a decision for people who never use it: another control to read past, another mode to misunderstand. Progressive disclosure — showing the primary options first and deferring secondary ones, a pattern Jakob Nielsen described in 2006 — helps, but it does not make clutter free.

Can we keep maintaining it? A capability is a long-term commitment: to its inputs, its providers, its documentation and its failure modes. Breadth multiplies that commitment.

The same discipline applies to research. Not every promising idea becomes a product, and some are better published than built; see Why Not Everything Ethen Researches Needs to Become a Product.

A worked example

The following example is illustrative.

Suppose a new model can generate short music clips, and someone proposes adding a "Make music" button to Chat.

The feature-count answer is easy: add the button, list the model, announce it. Within a week, Chat has another mode.

The depth answer starts with the questions. What work does it finish? Usually it is part of a creative project — a soundtrack for a video, a jingle for a product page — which means it needs references, versions and assets that outlive the conversation. Which surface owns it? Studio, where creative projects, history and spend already live. Can we make it dependable? Only once the route is qualified: supported inputs, validation, cost quoted up front, failure handling. Can we show evidence? A delivered clip should come with a record of what produced it. What does it cost Chat users who will never make music? Another button. Can we maintain it? Only if the provider and route are ones we are prepared to support.

The likely outcome is "not in Chat; in Studio, once the route is qualified." Chat might later offer a lightweight entry point that hands off to Studio. That is slower than adding a button. It is also how the capability ends up finishing work instead of decorating a menu.

Where breadth is the right answer

Choosing depth for workflows does not mean choosing narrowness everywhere. Some surfaces are valuable precisely because they are broad.

Reference surfaces. A model library is useful because it is comprehensive and organized. Ethen's catalog work grouped 1,499 provider endpoints into 491 model families so that breadth stays navigable; we explain why in Why Model Families Matter More Than Huge Model Counts. Breadth with structure is a feature of reference material.

Research. A research lab should explore more questions than products will ever answer. Breadth of inquiry is how you find the few ideas worth building.

Interoperability. Supporting open protocols and multiple model providers is breadth in service of choice. It is worth having when each connection is actually maintained.

The rule we try to follow is simple: be broad where people are looking things up, and deep where people are getting work done.

Tradeoffs and limitations

Depth is slower. Some capabilities that other products expose immediately will reach Ethen later, or only in one app. Some people will reasonably prefer a product that lets them try everything now.

Depth can become an excuse. "Not dependable yet" can turn into "never." We try to name what would change the answer, and to revisit it.

Boundaries drift. As the Chat picker shows, a focused design and a focused product are not the same thing. Keeping surfaces deep requires continual pruning.

We cannot prove the payoff yet. Depth is a product philosophy informed by research on feature fatigue and choice, not a measured Ethen result. We expect it to show up as work that finishes and people who come back; we will report evidence when we have it.

Some breadth is necessary to learn. Without trying things, you do not discover which deserve depth. Research previews and lightweight entry points are how we try things without committing every surface to them.

FAQ

Is it better for an AI product to have more features or fewer? It is better for each workflow to be complete. A product with fewer, deeper workflows usually finishes more work than one with many shallow capabilities, and it is easier to learn and trust.

Why do AI apps feel cluttered? Because exposing a new model capability is cheap and making it dependable is expensive. Products that add a mode or button for each model release accumulate options faster than they accumulate finished workflows.

What does product depth mean in AI? In Ethen's usage, depth means a workflow has coverage of the real task, reliability on supported inputs, recovery from failure, evidence of what was done, continuity across sessions and control over what it touches.

How does Ethen decide which features to build? By asking what work a capability finishes, which surface owns it, whether it can be made dependable, whether evidence can be shown, what it costs people who never use it, and whether it can be maintained.

Does depth mean Ethen will have fewer capabilities? Per surface, often yes. Across the family of apps, no: capabilities live in the app that owns them, and lighter surfaces hand off to deeper ones.

References

  1. Thompson, D. V., Hamilton, R. W., & Rust, R. T. (2005). Feature fatigue: When product capabilities become too much of a good thing. Journal of Marketing Research, 42(4), 431–442. https://doi.org/10.1509/jmkr.2005.42.4.431
  2. Iyengar, S. S., & Lepper, M. R. (2000). When choice is demotivating: Can one desire too much of a good thing? Journal of Personality and Social Psychology, 79(6), 995–1006. https://doi.org/10.1037/0022-3514.79.6.995
  3. Nielsen, J. (2006). Progressive disclosure. Nielsen Norman Group. https://www.nngroup.com/articles/progressive-disclosure/