Skip to main content

How do you version an AI agent so you can roll it back?

An AI agent is more than its prompt. Version the prompt, model, tools, policies and retrieval settings together as one release, record which release handled every request, and make rollback a configuration change rather than a redeploy.

A man in a dark suit holds a clipboard and checks a laptop at a standing desk while facing a large navy screen with a glowing cyan neural brain beside an AI chip.

Treat everything that shapes an agent's behaviour as one release: the prompt, the model and its settings, the tool definitions, approval policies and retrieval rules. Give each release a fixed version, record which version handled every request, and point production at a version by configuration, so rolling back means changing a pointer rather than shipping code.

Why agents need their own versioning

Traditional software changes when code changes, and version control already covers that. An AI agent can change behaviour without a single line of application code moving. Someone edits the system prompt, a tool description is reworded, a retrieval setting is tuned or the model provider ships an update. Each of these can alter what the agent says and does, and none of them shows up in a normal code release.

When something goes wrong after such a change, the team needs to answer two questions quickly: what changed, and how do we get back to the last version that worked? Without deliberate versioning, both answers depend on someone remembering what they edited. That is not a recovery plan.

Decide what a version contains

A version is only useful for rollback if it captures everything that affects behaviour. Rolling back the prompt alone does not help if the problem came from a new tool definition.

The parts that belong in a release

Include the system prompt and any prompt templates, the model name and its pinned version, sampling settings, the full set of tool definitions and descriptions, approval and permission policies, retrieval and routing rules, and any guardrail configuration. Store these together as one versioned bundle, outside the application code, with an identifier that never changes once published.

What stays outside the version

Data the agent reads, such as documents in a knowledge base, changes on its own schedule and is usually versioned separately. Note which index or content snapshot a release was tested against, so a change in data can be told apart from a change in the agent. Keeping that knowledge current is its own discipline, with its own owners and review cycle.

Pin production to an explicit version

Production should never run "whatever is latest". Each environment, from development through staging to production, should point at a named version, and moving a version forward should be a reviewed change. Some managed agent platforms now build this in: Anthropic's tutorial on prompt versioning and rollback for managed agents shows each update producing an immutable version, sessions pinned to a specific version, and rollback done by pointing callers back at the earlier one.

The same approach works on any stack. The review gate moves from "who edited the prompt" to "who promoted this version to production", which is the decision that actually matters.

Gate every promotion on evaluation

A new version should earn its place before it handles real work. Run it against the agent's evaluation suite and compare the results with the version currently in production. Look for regressions as well as improvements, particularly in the cases that have caused trouble before. The kinds of failure a good suite surfaces are covered in what AI agent evals catch.

Once a version passes, release it gradually rather than all at once. Send a small share of traffic to it, watch quality, cost and error signals, and widen the rollout only while those hold. This is the same staged pattern described in how to roll out an AI agent in stages, applied to every change rather than only to the first launch.

Make rollback a pointer change

If rollback requires a code change, a build and a deployment, it will be slow at exactly the moment speed matters. Design it so that returning to the previous version is a single configuration change that takes effect for new requests straight away. Keep earlier versions available and runnable, rather than deleting them when a new one ships.

Think through what happens to work already in progress. An agent partway through a long task may have started under one version and would finish under another. Decide in advance whether in-flight work completes on its original version or stops and resumes, especially for tasks that run over days.

Record the version on every request

You cannot roll back what you cannot trace. Every request the agent handles should log the exact version that produced it, alongside the model, settings and tools in use. This lets the team see when a problem started, which version caused it and which users were affected. It also turns the agent's history into evidence for reviews and audits, as set out in what an AI agent audit trail contains, and gives monitoring something specific to compare, as covered in how to monitor an AI agent in production.

CodeDTX's AI agent development and enterprise AI integration teams build release and rollback into agents from the start. If changes to your agent are hard to trace or undo, talk to CodeDTX about putting versioning in place.

Frequently asked questions

Is keeping prompts in Git enough to version an AI agent?

It is a useful start, but rarely enough on its own. Git tracks the text of a prompt, yet an agent's behaviour also depends on the model version, settings, tool definitions and policies, which often live elsewhere. Unless all of them are captured together and production is pinned to that combined release, you can restore a prompt and still be running a different agent from the one that last worked.

What happens when the model provider updates the model?

If production calls a model alias that the provider moves to a newer version, your agent changes without any action from your team. Pin a specific model version wherever the provider allows it, and treat moving to a new model as a release like any other, with evaluation and a staged rollout. When a pinned model is due to be retired, plan the move early, as covered in our post on model deprecation.

Who should be allowed to promote a new agent version?

Promotion to production should be a reviewed step, owned by the team accountable for the agent in production rather than by whoever edited the prompt. A product or domain owner should sign off on behaviour changes, and an engineer should confirm the evaluation results and rollout plan. Recording who approved each promotion gives a clear line of accountability if a version later has to be rolled back.

How do we roll back an agent that has already taken actions?

Rolling back the version stops further mistakes, but it does not undo actions already taken, such as records changed or messages sent. Those need their own recovery, which is why actions that cannot be reversed should sit behind approval steps. The request log, with the version recorded on each action, shows exactly what the faulty version did, so the team can review and correct each affected record.

Share this post

Contact us to build the right product

Talk to our engineers about your application, the systems it connects to, and what you want to build next.

Get in touch
Two people discussing work with a laptop