Research

The Science of Agentic ERP: Agents, Reasoning, Memory and Control

A research overview of the foundations of agentic enterprise systems: agent theory, reasoning, grounding, memory, safe action, autonomy and measurement.

12 min readeVitalyst Research, Teksalah
Abstract

Enterprise resource planning (ERP) systems were designed as systems of record: they capture transactions faithfully and leave interpretation and action to people. Large language models make a different design possible, in which software agents read the state of the enterprise, reason about it, propose or take actions and learn from the results. This article sets out the scientific foundations of such agentic ERP systems: the classical theory of intelligent agents, the role and limits of language models as reasoning engines, how agents are grounded in relational business data, how memory and vector search change their cost and behaviour, how actions can be taken safely in a transactional system, how autonomy should be graded and governed, and how the result can be measured. It closes with the open problems that remain.

1From systems of record to systems of action

Planning software for business began with material requirements planning (MRP) in the 1960s and 1970s, formalised by Orlicky (1975): given a master production schedule, bills of materials and stock, compute what to make and buy, and when. MRP II extended the calculation to capacity and finance, and in the 1990s enterprise resource planning integrated finance, supply chain, manufacturing and human resources around a single database (Davenport, 1998).

Across these generations the division of labour stayed the same. The system calculates and records; people interpret the output, decide and act, and then enter the consequences back into the system. ERP became the authoritative system of recordwhile judgement remained outside it. Analytics and dashboards shortened the path from data to insight but did not change the division: a dashboard still waits for someone to look at it.

Agentic ERP moves part of the interpretation and action into the system itself. The ERP becomes a system of action: software agents observe the state of the enterprise, reason about what it means, propose or take steps, and record what they did. The scientific questions this raises are old ones from artificial intelligence and human factors, now applied to the most consequential data a company has.

2What an agent is

In artificial intelligence an agent is anything that perceives its environment through sensors and acts on it through actuators (Russell and Norvig, 2021). Wooldridge and Jennings (1995) gave the definition most used since: a system is an intelligent agent when it is autonomous (it operates without direct intervention), reactive (it perceives its environment and responds in time), proactive (it pursues goals rather than only reacting) and social (it interacts with people and other agents).

The belief-desire-intention (BDI) model (Rao and Georgeff, 1995) adds a structure for practical reasoning that maps well onto business software. Beliefs are the agent's view of the world, which in an ERP is the state of orders, stock, ledgers and documents. Desires are its goals, such as keeping stock above safety levels or collecting receivables on time. Intentions are the plans it has committed to, such as a set of proposed purchase requisitions. The model makes explicit that an agent must revise its plans when its beliefs change, which in an ERP happens with every posted transaction.

Herbert Simon's account of bounded rationality (Simon, 1955) explains why agents are useful at all in management. Decision makers cannot attend to every signal; they satisfice within limits of attention and information. Agents that continuously watch the transactional record and surface what matters extend those limits, rather than replacing the judgement exercised within them.

3The agent loop in an enterprise

An agent in an ERP runs a loop that can be described in six steps:

  1. PerceiveRead events and state: new orders, receipts, postings, documents, questions from users.
  2. InterpretRelate them to goals and context: is this normal, urgent, risky?
  3. PlanDecide what to do: answer, recommend, prepare a transaction, escalate.
  4. ActExecute through the application: create a draft, route an approval, notify.
  5. ObserveSee the result: was it approved, rejected, corrected?
  6. LearnStore what worked, so the next similar case is handled better.
Figure 1. The agent loop in an enterprise system.

Two properties distinguish an enterprise agent from a general assistant. Its environment is a structured, transactional system with strict rules of consistency, and its actions have financial and legal consequences. Both shape every layer of the architecture described below.

4Language models as the reasoning engine

Large language models are trained on very large text corpora and can be adapted to a wide range of tasks; Bommasani et al. (2021) describe them as foundation models. Three findings make them usable as the reasoning component of an agent. Prompting a model to produce intermediate reasoning steps improves performance on multi-step problems (Wei et al., 2022). Interleaving reasoning with actions, and feeding the results of those actions back into the reasoning, lets a model plan and correct itself: the ReAct pattern (Yao et al., 2023). And models can learn when and how to call external tools such as calculators, search or databases (Schick et al., 2023).

The limits are equally important. Language models can produce fluent statements that are not supported by their inputs, a failure known as hallucination (Ji et al., 2023). Their outputs vary between runs, and their confidence is not a reliable guide to their accuracy. For business software the conclusion is architectural: a language model should never be the source of a business fact. Facts come from the system of record; the model's role is to interpret the request, decide which facts are needed, retrieve them through controlled tools and explain the result.

Reflexion (Shinn et al., 2023) shows that agents improve when they record feedback on their own attempts and use it in later ones. In an ERP the feedback is unusually rich: every approval, rejection and correction by a user is an explicit signal about the quality of a proposal.

5Grounding agents in enterprise data

Most enterprise facts live in relational databases, so the central technical problem of an ERP agent is translating a question in natural language into a correct query over the database. Text-to-SQL has been studied for decades and is now measured on public benchmarks such as Spider (Yu et al., 2018) and BIRD (Li et al., 2023). Accuracy on these benchmarks has risen sharply with language models, but BIRD was built precisely to show the gap that remains on large, messy, real-world databases, where knowing the meaning of the data matters as much as knowing SQL.

ERP databases make the problem harder and easier at once. Harder, because they contain hundreds or thousands of tables, with names that encode decades of design decisions, and because business meaning depends on document status, accounting dimensions and organisational context. Easier, because many ERP systems are metadata-driven: an application dictionary describes every table, column, reference, window and validation in machine-readable form. That dictionary is, in effect, a semantic map of the business that an agent can consult to choose the right tables, joins and filters.

Retrieval-augmented generation (Lewis et al., 2020) supplies the remaining context: instead of relying on what a model memorised, the system retrieves relevant definitions, policies, examples of correct queries and prior answers, and gives them to the model with the question. In an enterprise setting retrieval must respect the same access rules as the data itself, so that an agent never retrieves what the user could not see.

6Memory, embeddings and vector search

Agents need memory at several time scales. Working memory holds the current task. Episodic memory records past interactions: what was asked, how it was answered, whether the answer was accepted. Semantic memory holds general knowledge about the business, such as definitions of KPIs or the structure of the chart of accounts. Procedural memory holds how to do things, such as the sequence of steps to prepare a month-end accrual. The generative agents of Park et al. (2023) showed how a memory stream with retrieval by relevance, recency and importance produces coherent behaviour over time.

Memory is usually implemented with embeddings: numerical vectors that place texts with similar meaning close together. Finding the most similar past item becomes a nearest-neighbour search in a high-dimensional space. Exact search is too slow at scale, so systems use approximate methods, of which hierarchical navigable small world graphs (Malkov and Yashunin, 2020) are among the most widely used. Vector extensions now bring this capability into general-purpose relational databases, so that an ERP can keep its memory next to its transactional data, under the same backup, security and residency arrangements.

Memory changes the economics and the behaviour of an agent. A question semantically close to one already answered can reuse the earlier interpretation, query and context instead of starting again, which reduces calls to the language model and makes answers more consistent. The risks are equally concrete: a remembered result can become stale as transactions are posted, a remembered interpretation can be wrong and be repeated, and memory can leak information across users if it is not partitioned by permission. Sound designs reuse reasoning freely, revalidate facts against current data, and partition memory by tenant and role.

7Acting safely in a transactional system

An ERP is defined by its invariants. Every journal balances; documents move through defined states (draft, in progress, completed, closed, reversed); closed periods cannot be changed; quantities on hand reconcile with movements; segregation of duties separates who creates, who approves and who pays. These invariants are enforced by the application layer, through business rules, validations and workflow, not by the database alone.

It follows that an agent should act through the application, never around it. An agent that writes directly to database tables bypasses the rules that make the ledger trustworthy. An agent that calls the same services a user would call inherits every validation, every approval workflow and every audit record. Four further principles follow from the transactional setting:

  • Read by default. Answering questions requires only read access; write access is granted per action, not per agent.
  • Draft before commit. Agents create documents in a draft or proposed state; completing them is a separate, authorised step.
  • Reversibility. Every action has a defined reversal, using the ERP's own reversal documents rather than deletion.
  • Idempotency. Repeating an action because of a retry or a timeout must not create a duplicate transaction.

8Levels of autonomy and human control

Human factors research has long treated automation as a matter of degree. Sheridan and Verplank (1978) described ten levels between full human control and full automation, and Parasuraman, Sheridan and Wickens (2000) showed that the level can differ for each stage of a task: acquiring information, analysing it, deciding and acting. Their central lesson for ERP is that high automation of information acquisition and analysis is usually beneficial, while automation of decisions and actions should rise only as reliability is demonstrated and as the cost of error allows.

For enterprise systems a practical scale has six levels:

LevelNameWhat the system doesHuman role
0RecordStores transactionsInterprets, decides and acts
1AnswerAnswers questions from live dataAsks, decides and acts
2AdviseFlags issues and recommends in contextDecides and acts
3PrepareDrafts transactions and plansReviews and approves (human in the loop)
4Act within policyExecutes routine actions inside set limitsSets policy, handles exceptions (human on the loop)
5AutonomousActs without routine oversightAudits outcomes
Table 1. Levels of autonomy in an enterprise system, after Sheridan and Verplank (1978) and Parasuraman et al. (2000).

Different processes sit at different levels at the same time. A company may let an agent act within policy when matching bank statement lines, while purchase orders above a threshold remain at level 3 indefinitely. The level is a governance decision per process, not a property of the software.

9Multiple agents and tool protocols

A single general agent becomes unreliable as the number of tools and responsibilities grows. Systems therefore divide work among specialised agents, each with a narrow set of tools and goals (a planning agent, a payables agent, an analysis agent), coordinated by a router or orchestrator. Wang et al. (2024) survey the resulting architectures for profile, memory, planning and action.

Standard protocols are emerging for connecting agents to tools and data. The Model Context Protocol (Anthropic, 2024) defines how an AI application discovers and calls tools exposed by another system. For an ERP this means its capabilities can be offered to many AI applications through one controlled interface, with authentication, permission checks and logging applied at that interface rather than in each application.

10Measuring an agentic ERP

Claims about agentic systems are only as good as their measurement. The following measures are specific enough to be tracked in production:

  • Answer accuracy: the share of questions whose answer matches a verified result; for database questions, execution accuracy, as used in the BIRD benchmark.
  • Task success rate: the share of agent proposals accepted without change.
  • Intervention rate: the share corrected or rejected by a person, by process and by agent; a falling rate is the evidence for raising autonomy.
  • Error severity: not only how often an agent is wrong but what the error would have cost had it not been caught.
  • Time to decision: elapsed time from an event to the corresponding action, before and after the agent.
  • Cost per task: model usage, infrastructure and review time per completed task.
  • Memory reuse rate: the share of requests served with the help of local memory, and the accuracy of those reused answers.

11The economics of reasoning

Language model usage is priced per token processed, so the running cost of an agentic system can be written simply. For a given process:

cost per task = calls per task × (input tokens + output tokens) per call × price per token  +  review time × cost of reviewer time

Each term suggests a lever. Fewer calls per task come from memory that reuses prior reasoning and from routing simple requests to deterministic code. Fewer tokens per call come from retrieving only the relevant schema and context rather than everything. Lower price per token comes from using smaller models where they suffice. Lower review time comes from better proposals and from presenting them so that a reviewer can approve with confidence. The first lever, reuse through memory, also improves consistency, which is why memory is an economic and a quality decision at the same time.

12Governance and risk

The NIST AI Risk Management Framework (NIST, 2023) organises AI governance into four functions, Govern, Map, Measure and Manage, which translate directly to agentic ERP. Govern sets who may grant an agent which permissions and at what level of autonomy. Map records which processes agents touch and what could go wrong in each. Measure applies the metrics above continuously. Manage acts on them: lowering an agent's autonomy, retraining it on corrected examples or withdrawing it.

Three controls are specific to enterprise data. Every agent action must leave an audit record that identifies the agent, the user on whose behalf it acted, the data it used and the result, so that an auditor can reconstruct any decision. Access must be evaluated for the user, not for the agent, so that the agent cannot become a route around permissions. And data residency and confidentiality obligations apply to everything sent to a language model, which favours designs that keep memory and retrieval inside the company's own environment and send models only what each task needs.

13Open problems

  • Verification. Checking that an agent's proposal is correct can cost as much as producing it. Methods that let a reviewer verify quickly, by showing evidence, differences from normal and the rule applied, are as important as better agents.
  • Long-horizon planning. Agents handle single tasks well and multi-week plans less well; production and cash planning still need classical optimisation alongside language models.
  • Schema and process drift. Enterprises change their structures, codes and processes; agents and their memories must notice and adapt.
  • Evaluation on private data. Public benchmarks do not reflect any particular company; each deployment needs its own test set of questions and expected answers.
  • Accountability. When an agent acts within policy, responsibility rests with whoever set the policy. Organisations need to make that explicit before they raise autonomy.

14Conclusion

Agentic ERP is not a new kind of database or a chat window on an old one. It is a change in where interpretation and action happen: from people working around the system of record to agents working inside it, through its rules, under human governance. The science needed to build it well draws on decades of work on agents, human-automation interaction and databases, combined with recent advances in language models. Its success will be decided less by the intelligence of the models than by the quality of grounding, memory, control and measurement around them.

References

  1. Anthropic. (2024). Model Context Protocol specification. modelcontextprotocol.io.
  2. Bommasani, R., et al. (2021). On the opportunities and risks of foundation models. arXiv:2108.07258.
  3. Davenport, T. H. (1998). Putting the enterprise into the enterprise system. Harvard Business Review76(4), 121-131.
  4. Ji, Z., et al. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys55(12).
  5. Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems 33 (NeurIPS 2020).
  6. Li, J., et al. (2023). Can LLM already serve as a database interface? A big bench for large-scale database grounded text-to-SQLs. NeurIPS 2023, Datasets and Benchmarks Track.
  7. Malkov, Y. A., and Yashunin, D. A. (2020). Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence42(4), 824-836.
  8. NIST. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1.
  9. Orlicky, J. (1975). Material Requirements Planning. McGraw-Hill.
  10. Parasuraman, R., Sheridan, T. B., and Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics, Part A30(3), 286-297.
  11. Park, J. S., et al. (2023). Generative agents: Interactive simulacra of human behavior. Proceedings of UIST 2023.
  12. Rao, A. S., and Georgeff, M. P. (1995). BDI agents: From theory to practice. Proceedings of the First International Conference on Multi-Agent Systems (ICMAS-95)312-319.
  13. Russell, S., and Norvig, P. (2021). Artificial Intelligence: A Modern Approach (4th ed.). Pearson.
  14. Schick, T., et al. (2023). Toolformer: Language models can teach themselves to use tools. NeurIPS 2023.
  15. Sheridan, T. B., and Verplank, W. L. (1978). Human and computer control of undersea teleoperators. MIT Man-Machine Systems Laboratory.
  16. Shinn, N., et al. (2023). Reflexion: Language agents with verbal reinforcement learning. NeurIPS 2023.
  17. Simon, H. A. (1955). A behavioral model of rational choice. The Quarterly Journal of Economics69(1), 99-118.
  18. Wang, L., et al. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science18(6).
  19. Wei, J., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. NeurIPS 2022.
  20. Wooldridge, M., and Jennings, N. R. (1995). Intelligent agents: Theory and practice. The Knowledge Engineering Review10(2), 115-152.
  21. Yao, S., et al. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations (ICLR 2023).
  22. Yu, T., et al. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. Proceedings of EMNLP 2018.
How to cite this article

eVitalyst Research. (2026). The Science of Agentic ERP: Agents, Reasoning, Memory and Control. Teksalah. https://evitalyst.com/agentic-erp/science/

This article is written as a neutral overview of the field. eVitalyst, an agentic ERP platform by Teksalah, applies these principles; see architecture and governance for how.

Put the science to work

See how these principles run in a live ERP.