Skip to content

EthenEthenEthen

Research Proposal · 2026-10-03 · Enterprise / Sovereign AI

Process Memory: Learning How Organizations Actually Get Work Done

Publication type
Research Proposal
Research program
Enterprise / Sovereign AI
Published
Authors
Ethen Research Lab
Reading time
11 min read

An organization's documents say how work should be done. Its completed work shows how it is done. Agents need both, and today they mostly get the first.

Cover image for "Process Memory: Learning How Organizations Actually Get Work Done". Decorative abstract motif; contains no data.

Abstract

Enterprise AI systems invest heavily in remembering what an organization knows: indexing documents, extracting entities, building knowledge graphs and retrieving relevant passages. Much of what makes work succeed inside an organization is not written down. Who actually approves exceptions, which system is checked before a refund, which steps are skipped for long-standing customers, where escalations go on Fridays: this knowledge lives in how work gets done. This proposal distinguishes document memory from process memory and proposes building process memory for AI agents by mining the event logs of completed, verified tasks into per-organization process models. The models serve two purposes: planning hints for new tasks, and baselines for detecting runs that deviate from the usual process. We ground the proposal in process mining, an established field that discovers process models from event logs, and in recent work on agents that learn reusable workflows from experience. We discuss the central risk, that agents will learn habits that should not be repeated, and propose filtering by verified outcome and policy conformance. We also describe an experiment comparing agents with and without process memory. The proposal is untested.

Two kinds of memory

Agent memory research has concentrated on retaining and retrieving content: facts from conversations, documents and records. Systems page information between context and external storage (Packer et al.), extract and consolidate salient facts (Chhikara et al.), or track when facts were valid in temporal knowledge graphs (Rasmussen et al.). Enterprise search does the same at organizational scale.

Content memory answers the question what does the organization know or say? It does not answer how is this kind of work actually done here? Figure 1 contrasts the two.

Matrix comparing document memory and process memory on: source (documents, messages, records versus completed task event logs and receipts); question answered (what does the organization know or say versus how is this kind of work actually done here); typical structure (passages, entities, embeddings versus process models with steps, branches, approvals, exceptions); freshness risk (stale content versus process drift); main failure (retrieving outdated policy versus imitating a habit that should not be repeated).

Figure 1. Document memory versus process memory. Both are useful. They answer different questions, come from different sources and fail in different ways. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis.

The second question matters because procedures in real organizations differ from their documentation. A refund policy may say that refunds over a threshold need finance approval. In practice, approvals may be delegated to a team lead during certain periods. Refunds for one product line may first require checking a separate billing system. Customers in one region may receive credits instead of refunds. An agent that knows only the written policy will fail, escalate unnecessarily or act in ways that surprise the people who supervise it. Organizations accumulate this practical knowledge by doing the work. It is specific to each organization, it changes slowly, and it is hard to transfer.

What process mining offers

Process mining is a mature field concerned exactly with discovering how work is done from records of doing it (van der Aalst). Its basic input is an event log: a set of cases, each a sequence of events with an activity name, a timestamp, an actor and attributes. Event logs have a standard interchange format, XES (IEEE 1849-2016). From such logs, process-discovery algorithms produce process models showing frequent paths, branches, loops and exceptions. Conformance checking compares observed cases against a reference model to find deviations. Enhancement adds performance and decision information to models.

Agent systems that record their work in structured form, such as the Work Receipt, naturally produce event logs. Each task is a case. Each step, tool call, approval and outcome is an event, with an actor that may be a person or an agent principal. The Outcome Warehouse stores these records with verified outcomes attached, which conventional process mining usually lacks. Process mining often knows that a case finished, rarely whether it finished well.

Agents that learn workflows

Agent research has begun to learn procedural knowledge from experience. Agent Workflow Memory induces reusable workflows from past trajectories and supplies them to agents on later tasks, improving web-navigation benchmark performance (Wang et al.). Skill-library approaches accumulate executable procedures (Wang et al., Voyager; Zheng et al.). These methods learn the agent's successful procedures. Process memory, as proposed here, learns the organization's procedures, including steps performed by people and other systems, from the full record of completed work.

Proposal

Figure 2 shows the pipeline.

Pipeline: work receipts of completed tasks; event log with case identifier, activity, timestamp, actor, system (XES-style); process discovery per task family; filter by verified outcome and policy conformance; process model with frequent paths, approvals and known exceptions; two uses: planning hints for new tasks and anomaly detection for runs that deviate from usual process. All within the tenant boundary.

Figure 2. Mining process memory from verified work. Completed tasks are converted into an event log, mined into process models per task family, filtered by verified outcome and conformance to policy, and then used in two ways: as planning hints and as baselines for anomaly detection. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen research proposal; process-mining concepts.

1. Event log. Completed tasks are converted into an event log with case identifier, activity, timestamp, actor and system, plus the task's verified outcome and any policy evaluations.

2. Discovery per task family. Process models are discovered separately for each task family, such as refunds, access requests or vendor onboarding, since a single model across all work is unreadable and useless.

3. Filtering. Only cases with verified successful outcomes and conformance to current policy contribute to the model used for planning. This is the critical step, discussed below.

4. Use as planning hints. When an agent begins a task in a known family, the relevant process model is compiled into concise guidance: the usual sequence, who usually approves, which checks usually precede the consequential action, which exceptions are common. Guidance is advisory. It never grants authority beyond the agent's mandate.

5. Use for anomaly detection. A run that deviates sharply from the usual process, such as a refund with no preceding billing check or an approval from someone who never approves this kind of request, is flagged for review. Deviation is not proof of error, but it is a signal.

The central risk: learning bad habits

Mining "how work is actually done" will also learn shortcuts, workarounds and violations. If finance approvals are routinely skipped for small amounts in breach of policy, a naive process model will teach the agent to skip them too. Organizations also contain obsolete habits that persist only because nobody questioned them.

Three safeguards follow. Filter by verified outcome: only cases whose outcomes were verified as successful contribute to planning models, so processes that usually fail are not imitated. Filter by policy conformance: cases that violated current policy are excluded from planning models even if they succeeded, and are instead reported as compliance findings. Separate description from prescription: the full, unfiltered model is retained for analysis and anomaly detection. Only the filtered model guides agents. When the two differ, the difference is itself useful to the organization: it shows where actual practice departs from policy.

Keeping process memory current

Processes change. A new approver is appointed, a system is retired, a policy threshold moves. A process model mined from last year's work will confidently guide agents along paths that no longer exist. Three practices keep process memory current. Windowing: models are mined from a recent window of cases, with the window long enough to include rare exceptions and short enough to reflect current practice; the trade-off is set per family. Versioning: each model is versioned with the window and filters used, so that guidance given to an agent can be traced to the model that produced it, and a model can be withdrawn if it turns out to be wrong. Drift detection: conformance of new cases to the current model is monitored continuously, and a sustained rise in deviations triggers re-mining and human review. Sudden drift is often the first sign that something in the organization changed, which is valuable information in its own right.

Process memory also inherits the rights of the records it was mined from. If a set of cases is withdrawn from permitted use, models built from them must be rebuilt without them, following the lineage practices in Rights as Infrastructure. Because models summarize many cases, the effect of removing a few is usually small. That is a reason to make rebuilding cheap, not a reason to skip it.

Process memory in context

Process memory complements the evidence-preserving context compiler. The compiler decides how to fit relevant material into a bounded context; process memory supplies one kind of highly relevant material, a compact description of how this task is usually done. A process hint is small: a few lines summarizing a model, not a log of past cases.

Privacy, rights and boundaries

Process models are derived from an organization's work and describe its internal operations, approvers and exceptions. They are tenant-private by default, and are mined and stored inside the tenant boundary. Even aggregate process statistics can reveal sensitive information about how an organization operates. Sharing them across organizations requires explicit grants, and any cross-tenant learning about processes would need the protections discussed in Private AI Improvement Without Raw Data Export. Event logs also identify individual employees as actors. Their use for anything resembling performance monitoring raises employment-law and works-council questions that require legal review before deployment.

Experiment design

Research question. Does process memory, mined from verified and policy-conformant work, improve agents' verified outcomes or reduce human escalations and approvals on new tasks, compared with document memory alone?

Figure 3 shows the design.

Experiment: held-out tasks from families with mined process models. Three arms: no memory, document memory only, document plus process memory. Metrics: verified success, approvals and escalations per task, policy violations, cost per verified outcome, anomaly detection precision and recall on seeded deviations. A note warns that copying habitual but non-compliant processes counts as a failure.

Figure 3. Experiment design for process memory. The same agent, tools and budget, with and without process memory, on held-out tasks from families with mined processes. Process memory must improve verified outcomes or reduce approvals and escalations, not merely make plans look more familiar. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

The experiment uses held-out tasks from families that have enough historical cases to mine, with three arms: no memory, document memory only, and document plus process memory. Model, tools and budget are fixed across arms. The primary metrics are verified success, approvals and escalations per task, and policy violations. Cost per verified outcome is secondary. A separate evaluation seeds deviations into held-out runs and measures anomaly-detection precision and recall. Tasks can be generated in a synthetic enterprise with specified process models, which allows the ground truth of "how work is done here" to be known exactly. Real tenants are needed for the decisive test, along with enterprise web environments where agents already operate on realistic platforms (Drouin et al.); in-tenant evaluation follows the approach in Tenant Replay.

Falsifiers. Process memory is not worth building if it does not improve verified outcomes or reduce escalations relative to document memory, if it increases policy violations by teaching habits, or if mined models are too unstable over time to be useful.

An illustrative case

[ILLUSTRATIVE EXAMPLE — not an Ethen result.] In a simulated company, refund requests for annual subscribers are handled differently from monthly ones: support confirms the renewal date in the billing system, then a billing specialist, not the finance team, approves. The written policy mentions only finance approval. An agent with document memory routes approvals to finance, which bounces them back. An agent with process memory, mined from verified past refunds, checks the renewal date and routes to the billing specialist, matching practice. The mined model also shows that a quarter of small refunds skipped the billing check. Because those cases violate policy, they are excluded from the planning model and reported to the organization as a compliance finding rather than taught to the agent.

Limitations

Process memory requires enough historical cases per task family, which new deployments lack. Process discovery on agent-produced logs has not been evaluated, and logs that mix human and agent actions may be harder to mine. Filtering by policy conformance requires policies expressed precisely enough to check. Processes change, so models must be refreshed and drift monitored. Most importantly, the claim that process memory improves outcomes beyond document memory is a hypothesis.

Conclusion

Organizations know how to do their work in ways that their documents do not record. Agents that do that work need access to that knowledge, filtered so that they learn what works and complies rather than what merely happens. Process mining has the methods; agent systems that record verified outcomes have the data. The proposal is to join them, carefully, inside each organization's boundary.

FAQ

What is process memory? A per-organization model of how a kind of work is actually performed, mined from records of completed tasks, as distinct from memory of documents and facts.

Won't agents copy bad practices? They would if models were mined naively. The proposal uses only cases with verified successful outcomes that conform to current policy for guidance, and reports non-conforming practice separately.

Does process memory leave the organization? No. It is mined and stored inside the tenant boundary by default.

References

  1. van der Aalst, W. (2016). Process Mining: Data Science in Action (2nd ed.). Springer. https://doi.org/10.1007/978-3-662-49851-4
  2. IEEE (2016). IEEE 1849-2016: Standard for eXtensible Event Stream (XES) for Achieving Interoperability in Event Logs and Event Streams. https://doi.org/10.1109/IEEESTD.2016.7740858
  3. Wang, Z. Z. et al. (2024). Agent Workflow Memory. arXiv:2409.07429. https://arxiv.org/abs/2409.07429
  4. Wang, G. et al. (2023). Voyager. arXiv:2305.16291. https://arxiv.org/abs/2305.16291
  5. Zheng, B. et al. (2025). SkillWeaver. arXiv:2504.07079. https://arxiv.org/abs/2504.07079
  6. Packer, C. et al. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560. https://arxiv.org/abs/2310.08560
  7. Chhikara, P. et al. (2025). Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. arXiv:2504.19413. https://arxiv.org/abs/2504.19413
  8. Rasmussen, P. et al. (2025). Zep: A Temporal Knowledge Graph Architecture for Agent Memory. arXiv:2501.13956. https://arxiv.org/abs/2501.13956
  9. Drouin, A. et al. (2024). WorkArena. arXiv:2403.07718. https://arxiv.org/abs/2403.07718

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Product

    How We’re Rethinking AI Memory Across Ethen

    AI memory should not be one opaque pile of things an assistant decided to remember. Ethen's direction for memory rests on five ideas. Different kinds of memory get different rules: your preferences, facts about you, a project's context, your organization's knowledge, learned procedures and commitments you made each have their own scope, lifetime and controls. Memory is not the record of what happened: the authoritative record of a task's actions and outcomes is kept separately and never replaced by a summary. Memory carries its source and its validity, so it can be checked, superseded and explained. Permissions travel with memory, so a summary of a restricted document stays restricted and revoking access revokes what was derived from it. And you can see, correct, export and delete what Ethen remembers, and see when a memory influenced what Ethen did. This article explains each idea and what it means across Ethen's apps. It describes direction, not shipped architecture.

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.