Skip to content

EthenEthenEthen

Research Proposal · 2026-10-03 · Context / Skills / Transfer

Skill IR: Toward Model-Independent Agent Capabilities

Publication type
Research Proposal
Research program
Context / Skills / Transfer
Published
Authors
Ethen Research Lab
Reading time
13 min read

Most agent "skills" are prompts tuned for one model. When the model changes, the skill silently degrades. We propose separating what a skill promises from how a particular model realizes it.

Cover image for "Skill IR: Toward Model-Independent Agent Capabilities". Decorative abstract motif; contains no data.

Abstract

Agent systems accumulate AI agent skills: procedures for recurring work such as reconciling invoices, triaging a bug or drafting a contract amendment. Today these are usually stored as prompt text, tool descriptions or example trajectories, tuned for whichever model was current when they were written. When a better model arrives, they may continue to work, work worse, or fail in ways nobody notices. This proposal introduces Skill IR, an intermediate representation for skills that separates three parts with different lifetimes. The contract states inputs, outputs, preconditions, required tools, declared permissions and effects, failure modes and an attached verifier, and is meant to stay stable across models. Realizations are the model-specific instructions, examples or code that implement the contract, and are expected to change. Evidence records where the skill has been validated and where it has failed. By analogy with compilers, a contract is "lowered" into realizations for each target model family, and every realization must pass the same tests. We relate the proposal to research on skill libraries and workflow memory, and we describe how Skill IR would be validated, versioned and secured. It is an Ethen research proposal; no Skill IR system has been built.

The problem: skills bound to one model

Organizations that use agents for recurring work quickly build up procedural knowledge: how to handle this kind of request, in this system, under these rules. Capturing that knowledge as reusable skills is one of the clearest ways an agent system can improve without retraining a model. The research literature supports the idea. Voyager accumulated an executable library of skills in an open-ended environment and reused them on new tasks (Wang et al.). Agent Workflow Memory induces reusable workflows from past trajectories and improves performance on web-navigation benchmarks (Wang et al.). SkillWeaver lets web agents discover skills, refine them into reusable programmatic interfaces, and share them with other agents (Zheng et al.). ExpeL extracts insights from experience that later tasks can use (Zhao et al.).

These results share a weakness that matters in deployment: a skill is usually expressed in a form tuned to the model that produced or uses it. Prompt behavior is sensitive to formatting choices that carry no meaning (Sclar et al.). Hosted model behavior changes between versions (Chen et al.). Prompts optimized automatically for one model are not guaranteed to be optimal, or even adequate, for another (Yang et al.). A skill library built over a year may be partly obsolete after a single model upgrade, and nobody will know which parts until tasks start failing.

There is also a governance gap. A skill that issues refunds should declare that it moves money, which permissions it needs and how its success is checked. A prompt recipe declares none of these things. Skills packaged as instruction folders or tool descriptions do better, but rarely carry their verifier or their validation history.

Proposal: separate contract, realization and evidence

A Skill IR record has three parts (Figure 1).

Three stacked groups. Contract (stable): identity and version, typed inputs and outputs, preconditions and state assumptions, required tools, declared permissions and effects, failure modes, attached verifier. Realizations (model-specific, replaceable): instructions and examples for model family A, tool configuration for family B, code implementation where deterministic. Evidence (append-only): contract-test results, held-out promotion results, compatibility records by model and tool version, known negative transfer.

Figure 1. Anatomy of a Skill IR record. A skill is split into three parts with different lifetimes. The contract states what the skill does and must stay stable across models; realizations are model-specific and expected to change; evidence records where the skill has actually been validated. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen research proposal (Skill IR).

The contract describes what the skill does, independent of any model:

  • identity and semantic version;
  • typed inputs and outputs;
  • preconditions and state assumptions ("the invoice exists and is unpaid");
  • required tools, by interface rather than implementation;
  • declared permissions and effects ("reads billing records; may issue a refund up to the mandate's limit");
  • known failure modes ("duplicate invoice numbers across subsidiaries");
  • an attached verifier that decides whether an invocation succeeded.

Realizations implement the contract for particular targets: instructions and examples tuned for one model family, tool configurations for another, or deterministic code where the behavior does not need a model at all. Realizations are expected to change and to be replaced.

Evidence is an append-only record of validation: contract-test results, held-out promotion results, and compatibility records by model and tool version, including known negative transfer, meaning versions where the skill performs worse than having no skill.

The contract is the stable artifact. It is what a mandate can reason about ("this skill moves money, so it needs approval above the threshold"). It is what an auditor can read, and what carries across model generations.

Lowering, by analogy with compilers

Compilers translate source programs into an intermediate representation, then lower that representation to machine code for specific targets. The IR is where target-independent meaning lives; target-specific detail is added late. Skill IR borrows the idea (Figure 2).

One skill contract at left, arrows to three realizations for model families A, B and C. Each realization runs the same contract tests and held-out evaluation. Results feed a compatibility table: family A validated, family B validated with lower success, family C failed and marked incompatible.

Figure 2. Lowering one contract to several model families. By analogy with compiler intermediate representations, one capability contract is lowered into realizations for each target model family. Every realization must pass the same contract tests and held-out evaluation before its compatibility record is marked valid. Evidence label: ILLUSTRATIVE — NOT MEASURED ETHEN DATA. Source: Ethen research proposal; outcomes shown are illustrative.

A single contract is lowered into realizations for each model family. The realizations may differ in wording, in the number and form of examples, in tool-call conventions, or in whether a step is handled by the model or by code. Every realization must pass the same contract tests and the same held-out evaluation before its compatibility record is marked valid. When a new model family arrives, the contract does not change. A new realization is produced, by a person, by an optimizer or by the new model itself, and tested against the contract.

The analogy has limits. Compilers lower programs with precise semantics; skill contracts describe behavior that models realize probabilistically. Lowering is therefore a search for a realization that passes tests, not a deterministic translation, and validation carries the weight that correctness proofs carry in compilers.

How Skill IR compares with existing representations

Figure 3 compares common skill packaging approaches against the properties Skill IR requires.

Matrix of five representations (prompt recipe, tool description, code function, induced workflow memory, Skill IR record) against six properties: declared preconditions, declared permissions and effects, attached verifier, explicit failure modes, portable across model families, versioned validation evidence. Skill IR is designed to have all six; prompt recipes have none reliably.

Figure 3. How skill representations compare. Common ways of packaging agent capabilities, assessed against the properties Skill IR requires. The comparison is qualitative and concerns the representations as typically used, not any particular implementation. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis; qualitative.

Models can learn when and how to call tools from their descriptions (Schick et al.), and tool descriptions, such as schemas exposed through the Model Context Protocol, are portable across models and typed, but they describe a single operation rather than a procedure, and they usually lack verifiers and failure modes. Code functions are portable and testable but cannot capture judgment-heavy steps. Induced workflow memories are close to what agents actually need, but they are rarely typed or permissioned. Skill IR does not replace any of these. A Skill IR realization may use tool schemas, code and workflow examples. It adds the contract and the evidence around them.

We do not propose a new proprietary packaging format. Interoperability matters more than ownership. A Skill IR record should be expressible on top of existing open standards for tools and skill packaging, with the contract and evidence as additional, openly documented metadata.

An illustrative contract

[ILLUSTRATIVE EXAMPLE — not an implemented Ethen skill.] Consider a skill for reconciling a vendor invoice against a purchase order. Its contract accepts an invoice identifier and returns one of three typed results: matched, mismatched with a list of discrepancies, or unable to determine with a reason. Its preconditions are that the invoice exists, is in an unpaid state, and references a purchase order the caller may read. It requires two tool interfaces, one to read invoices and one to read purchase orders, and declares read-only access: it may not approve, pay or modify anything. Its known failure modes include partial deliveries split across invoices, currency conversions with different rate dates, and duplicate invoice numbers across subsidiaries. Its verifier recomputes the match deterministically from the underlying records and compares it with the skill's result.

Nothing in that contract mentions a model, a prompt or an example. A realization for one model family might include detailed instructions and three worked examples; another might delegate the arithmetic to code and use the model only to interpret free-text line items. Both are judged by the same verifier, and both carry their own compatibility evidence.

Versioning and lifecycle

Skill IR needs versioning rules that distinguish kinds of change. A change to the contract, such as a new output type, a widened permission or a removed precondition, is a breaking change. It requires a new major version, a fresh review of permissions and a new evidence record. A new or modified realization that passes the existing contract tests is a compatible change. A skill's dependencies are tool interfaces, not tool implementations. When a tool's interface version changes, every skill that depends on it is flagged for revalidation, in the same way that a dependency upgrade triggers tests in software.

Skills also need deprecation. A skill whose realizations fail on every currently supported model family, or whose underlying tools have been retired, should be marked deprecated rather than silently removed. Its evidence remains as a record of what once worked and why it stopped.

Validation and promotion

A skill earns trust through evidence, not authorship. We propose three gates.

Contract tests. Executable checks that a realization respects the contract: it produces typed outputs, refuses when preconditions fail, never exceeds declared permissions, and is judged successful by the attached verifier on a set of reference cases.

Held-out promotion. A new or optimized skill version is promoted only if it improves verified outcomes on tasks held out from its development, without regressing existing skills. This matches the skill-optimization rung in Ethen's internal improvement ladder: skills may be optimized automatically, but only through held-out gates.

Revalidation on change. When a model or tool version changes, every affected skill is re-run against its contract tests and a held-out sample before its compatibility record is extended. The data structure for these records is the Capability Transfer Ledger. The experiment that asks whether skills survive a frontier-model upgrade is How to Measure Whether AI Skills Survive a Frontier-Model Upgrade. The benchmark track that measures transfer across models and tools is VerifiedWork Transfer.

Verifiers travel with skills

Attaching a verifier to every skill is the most consequential design choice in this proposal. It means that every invocation of a skill can be judged by the same standard regardless of which model realized it. That makes compatibility measurable, and it means a skill's output can enter verified experience without extra work. It also means the skill inherits its verifier's errors. A skill whose verifier accepts plausible but incorrect outputs will be promoted and spread. The reliability measurements in Evaluating the Evaluators therefore apply to skill verifiers as much as to any other.

Security: skills are a supply chain

Skills that can be shared across teams or organizations are a software supply chain. A malicious or careless skill can request permissions it does not need, exfiltrate data through tool calls, or hide instructions that redirect an agent. Several protections follow from the design:

  • Declared permissions are enforced, not advisory. A skill that declares read-only access cannot write, regardless of what its realization instructs.
  • Signed provenance. Skill records are signed by their authors, and their evidence by the systems that produced it. Supply-chain attestation frameworks such as in-toto show how signed claims about each step can be verified end to end (Torres-Arias et al.).
  • Untrusted text stays untrusted. Instructions inside a skill realization are treated as content, not policy; they cannot widen a mandate.
  • Private catalogs first. Sharing should begin within a tenant, then across consenting tenants, and only later publicly, with review.

Research questions

  1. Portability. Do typed contracts with attached verifiers transfer across model families better than prompt recipes or retrieved procedures, at matched effort?
  2. Revalidation cost. How much effort does it take to produce a passing realization for a new model family, compared with rewriting the skill from scratch?
  3. Negative transfer detection. Can contract tests detect, before deployment, the cases where a skill makes a new model worse?
  4. Optimization. Can skill realizations be optimized automatically against held-out gates without overfitting to the gate?

Risks and counterarguments

Better models may not need skills. A sufficiently capable model may perform the procedure from a short description. If so, the contract still has value as documentation, permission declaration and verifier binding, but the realization may shrink to almost nothing. That would be a good outcome.

Abstraction leakage. Some skills may depend on model-specific behavior so deeply that no stable contract captures them. The evidence records will reveal this as persistent incompatibility.

Maintenance burden. Contracts, tests and evidence for hundreds of skills cost effort. Skill IR is justified only if that effort is smaller than the cost of undetected skill degradation.

Limitations

This proposal has not been implemented or tested. The compiler analogy is motivating, not exact. The claim that contracts transfer better than prompt recipes is a hypothesis. Verifiers for judgment-heavy skills may be too unreliable to support the validation scheme. Security properties depend on enforcement infrastructure that is described here only at a high level.

Conclusion

An organization's skills are among its most valuable operational knowledge, and today they are written in the most perishable form available: text tuned for one model. Skill IR proposes to keep the durable part, what the skill promises, which permissions it needs and how success is checked, separate from the perishable part, how a given model is told to do it. Whether that separation actually preserves capability across model generations is the question the Transfer Ledger is designed to answer.

FAQ

Is Skill IR a new file format? It is a representation, a set of fields and validation rules, that can sit on top of existing tool and skill packaging standards. We do not propose replacing them.

Who writes realizations for a new model? People, automated optimizers, or the new model itself. In every case the realization must pass the contract tests and held-out evaluation.

Does every skill need a verifier? In this design, yes. Without one, compatibility cannot be measured and outputs cannot become verified experience.

References

  1. Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291. https://arxiv.org/abs/2305.16291
  2. Wang, Z. Z. et al. (2024). Agent Workflow Memory. arXiv:2409.07429. https://arxiv.org/abs/2409.07429
  3. Zheng, B. et al. (2025). SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills. arXiv:2504.07079. https://arxiv.org/abs/2504.07079
  4. Zhao, A. et al. (2023). ExpeL: LLM Agents Are Experiential Learners. arXiv:2308.10144. https://arxiv.org/abs/2308.10144
  5. Sclar, M. et al. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. arXiv:2310.11324. https://arxiv.org/abs/2310.11324
  6. Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
  7. Yang, C. et al. (2023). Large Language Models as Optimizers. arXiv:2309.03409. https://arxiv.org/abs/2309.03409
  8. Schick, T. et al. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv:2302.04761. https://arxiv.org/abs/2302.04761
  9. Model Context Protocol. Specification. https://modelcontextprotocol.io/specification
  10. Torres-Arias, S. et al. (2019). in-toto: Providing farm-to-table guarantees for bits and bytes. USENIX Security 2019. https://www.usenix.org/conference/usenixsecurity19/presentation/torres-arias

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Product

    Ethen Code: What We're Building Next

    Ethen Code is moving from assistive coding — explaining code, writing functions, fixing snippets — toward software tasks that end with evidence a person can check. The direction has seven parts: reproduce a problem before changing anything; plan a minimal, reviewable change before executing it; work in isolated environments with scoped permissions; treat tests, builds and review as evidence rather than as a finish line; put approvals in front of merges, deployments and other consequential steps; recover from failures in long-running work without repeating effects; and keep the same task model across Chat, the cloud workspace and Desktop. This article describes that direction. It is not a release schedule, and it makes no availability or date claims.

  • Security & Trust

    How Ethen Thinks About AI Actions That Can’t Be Undone

    Ethen treats AI actions that can't be undone — sending an external message, paying a third party, permanently deleting data, accepting legal terms, disclosing data outside where it is allowed to go — as a separate class of action with its own rules. The approach starts with one question asked before any action: if this is wrong, what does it take to make it right? Actions are classified on a reversibility scale. Where Ethen controls the system, it tries to move actions toward the reversible end with previews, staging and undo windows. For what remains irreversible, a person approves the exact action, the action is performed once with an identifier tied to its intent, and its effect is confirmed afterward. Powers that would undermine every other control — changing permissions, changing policy — are not delegated to agents at all.

  • Company

    From Screen to Physical World: How We Think About Ethen Robotics

    The Ethen Robotics Research Lab is a research direction, not a hardware program. It studies the problems every acting agent faces — pursuing goals over many steps, understanding the state of its environment, predicting what its actions will do, recovering when they go wrong, knowing when to stop, and proving that a task is done — and it studies them in software first. Its research questions run from near-term work on reliable action in software, through skills that survive changes in models and tools and digital state models that predict effects before acting, to the far-off question of grounding any of this in the physical world. Physical work, if it ever happens, would come through integration with existing systems, ground truth gathered with partners, and control only by qualified robotics and safety teams under recognized standards. No products, partnerships, dates or results are announced here. This article explains what the Lab studies and how we think about the path from screen to physical world.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.