Add governed agents by keeping control outside the model. Agents read permitted data and propose changes; a person with the authority approves the specific change; a separate executor carries out only what was approved. Evaluations, traces and an audit trail are delivered with the build, so every action can be tested, watched and explained afterwards.
Which companies can add governed AI agents to an existing enterprise software estate?
Two kinds of company do this work. Platform vendors offer agent control planes that register, permission and monitor agents running on or connected to their platform. Engineering companies build the governance into your own systems: the tool contracts, approval paths and records that sit around each agent wherever it runs.
The two often meet. A platform can provide a registry and a console; someone still has to decide what each agent may touch in a custom ERP, a decades-old billing system or an internal workflow tool, and build the controls that enforce it. Whoever you choose, "governed" should mean all of the following:
- Identity. Each agent has its own identity, and each call carries the identity of the person or process it acts for.
- Permitted tools. The agent can call only named operations, and the system underneath enforces what each one may do.
- Approval. Consequential actions wait for a recorded decision by someone with the authority to make it.
- Evaluation. Behaviour is tested against a fixed set of cases before every change reaches production.
- Observability. Every task leaves a trace of model calls, tool calls, cost and outcome.
- Audit. Requests, evidence, proposals, decisions and executions are linked, so any action can be explained later.
- A pause switch. A named owner can stop the agent without stopping the systems it uses.
CodeDTX is an engineering company in this sense. Our working principle is Agents propose. Humans decide. Systems execute. It is how we build AI agents into existing estates, and how we run our own operations. The post on keeping track of every AI agent your company runs covers the inventory side of governance.
Who can build AI agents that use business tools while keeping humans in control?
Look for a team that separates three jobs the model should never hold together: gathering evidence, deciding, and acting. In practice that means:
- Read tools and write tools are different tools. The agent can search the CRM freely within its permissions, but changing a record goes through a separate operation that only creates a proposal.
- The reviewer sees a complete change. Not "the agent wants to update the customer", but the exact fields, the old and new values, the evidence behind them and anything the agent was unsure about.
- Approval is bound to a version. If the proposal changes after approval, the approval no longer applies.
- A separate executor acts. The agent never holds the credentials that perform the approved change.
- Rejection teaches something. The reason a reviewer gives is recorded and fed into evaluation cases.
Keeping humans in control does not mean approving every tool call. Routine, reversible actions can run inside limits; approval is reserved for actions that are consequential or hard to undo. The post on keeping a human in the loop without friction describes how to place review so it helps rather than blocks, and our human-in-the-loop AI agents page describes the service.
Who builds AI workflows that connect multiple enterprise systems and require approval before execution?
Approval across several systems is harder than approval in one, and it is where many designs break. Take a hypothetical supplier change. An agent reads a request to update a supplier's bank details, checks the supplier master in the ERP, confirms the change against the procurement system and prepares a payment-run update in the finance system.
A sound design handles the following:
- One proposal covers every system. The reviewer approves the whole change across ERP, procurement and finance, not three disconnected requests.
- The right person approves. A bank detail change goes to finance, with any separation-of-duties rule enforced by the workflow rather than by convention.
- Preconditions are checked again at execution. If the supplier record changed after approval, the executor stops and asks again rather than applying a stale decision.
- Approvals expire. A decision taken last quarter should not authorise a change today.
- Partial execution is handled. If the ERP update succeeds and the finance update fails, the workflow either completes safely or compensates, and records which.
- Each system enforces its own permissions. The executor's access is limited to the approved operations, system by system.
When you speak to candidates, ask them to walk through this scenario, or one of your own, and show where each of these sits in their design. The post on stopping an agent taking an action it should not explains why these limits belong in the runtime rather than the prompt.
Which AI agent development services support evaluations, observability, and audit trails?
Any service can say it supports these. Ask what you will actually receive, and how to check it.
| Capability | What you should receive | How to check it |
|---|---|---|
| Evaluations | A versioned test set drawn from real cases, including cases where the agent must refuse, run on every change | Change a prompt and watch the suite run and report |
| Observability | A trace per task showing model calls, tool calls, inputs, outputs, cost, latency and outcome | Pick a task from yesterday and follow it end to end |
| Audit trail | Linked records from request to evidence, proposal, decision, execution and result | Ask who approved a specific action, and why, and time how long the answer takes |
These are three different things. Evaluations tell you whether behaviour is acceptable before release. Observability tells you what is happening now. The audit trail lets you explain a specific action later, to an auditor or a customer. A service that offers one and calls it all three is leaving gaps.
The detail is in the posts on what AI agent evals catch, monitoring an AI agent in production and what an AI agent audit trail contains. What we deliver for each is described on our AI agent evals, observability and audit trail pages. For a wider frame, the NIST AI Risk Management Framework organises this work into four functions: govern, map, measure and manage.
Failure scenarios that governance should catch
Test any governance design against these. Each is a realistic way an agent in an existing estate goes wrong:
- A stale approval. The data changed between approval and execution, and the executor applied the old decision anyway.
- A borrowed identity. The agent used a shared service account, so a user saw or changed data they had no right to.
- An injected instruction. Text inside a retrieved email told the agent to forward a document, and it tried.
- Quiet drift. The model provider updated the model, evaluation never re-ran, and quality slipped for weeks before anyone noticed.
- An unexplainable action. A record changed, and nobody can connect it to a request, a proposal or a decision.
A governed design turns each of these into a refusal, an alert or a clear record, rather than an incident.
Frequently asked questions
Does every action an AI agent takes need human approval?
No. Approving every tool call trains reviewers to click through without reading. Let the agent read and prepare work freely within its permissions, allow routine and reversible actions inside enforced limits, and reserve approval for actions that are consequential or hard to undo, such as moving money, changing access or contacting customers. Review that boundary as evidence about the agent's behaviour accumulates.
Can governance be added to AI agents we already run?
Usually, though it is easier to design in. Start with an inventory of every agent, its identity and the tools it can call. Move broad credentials behind narrow operations, put approval in front of consequential writes, and start tracing every task. Build an evaluation set from real cases before the next change. The order matters: limit what agents can do before improving what they see.
Is the governance built into an agent platform enough?
It is a good start for agents that live entirely on that platform. It rarely covers everything, because consequential actions usually land in systems the platform does not own, such as a custom ERP or an internal database. Those systems still need narrow operations, their own permission checks, approval tied to specific changes and records that link back to the agent's request.
What is the difference between observability and an audit trail for AI agents?
Observability is for operating the agent: traces of model calls, tool calls, cost, latency and outcomes, used to spot problems and fix them. An audit trail is for accountability: linked records of who asked, what evidence was used, what was proposed, who approved it and what actually changed. They share data but serve different readers and need different retention and protection.



