Skip to main content

How do you run an AI agent on a task that takes days?

Long agent tasks fail when they are run like long conversations. Treat them as durable workflows: save state after each step, make actions safe to repeat, pause properly for people and recheck facts before resuming.

A businessman and a humanoid robot work side by side at a desk by a window overlooking a city lit up at night.

Run an AI agent on a task that takes days by treating it as a durable workflow, not a long conversation. Save the state after every step, make each action safe to repeat, pause while waiting for people or systems, and recheck facts before resuming. The agent should survive a restart without losing work or repeating anything.

Why long tasks break agents built for chat

Most agents start life as a loop inside a single request: read the goal, call a tool, read the result, call the next tool, reply. That works when the whole task finishes while the user waits. It stops working when the task depends on things that take time, such as a manager's approval, a supplier's reply, an overnight batch job or a document someone has promised to upload.

Consider a hypothetical supplier onboarding agent. It collects the supplier's documents, requests a compliance check, waits for procurement to approve the terms, and then creates the vendor record in the finance system. Each step is simple. The difficulty is the gaps between them. During those gaps servers are redeployed, sessions expire, tokens are rotated, the model version changes and the facts the agent relied on move.

An agent that holds its progress only in memory or in a growing chat transcript loses everything the first time the process restarts. Worse, it may start again from the beginning and repeat work that already reached another system.

Treat the task as a workflow with saved state

The fix is to separate the task from the process that happens to be running it. Model the task as a workflow with named steps, and write its state to durable storage after every step: what has been done, what each tool returned, what the agent decided and what it is waiting for. Any worker should be able to pick up the task from that record and continue.

Keep that state structured rather than relying on the transcript. A record that says the compliance check was requested, with its reference and the time it was sent, is far more useful than a paragraph of conversation the model has to reread and interpret. Structured state is also cheaper to load, easier to inspect and harder for the model to misremember.

Workflow engines built for durable execution handle much of this for you: they persist each step, replay safely after a crash and let a run sleep until an event arrives. Whether you use one or build a simpler state table, the principle is the same. The model decides what to do next. The workflow remembers what has already happened.

Make every step safe to run twice

A long task will be interrupted, and some interruptions will happen in the middle of a step. When the worker comes back, it may not know whether its last action reached the other system. If the step simply runs again, the supplier might be created twice or the same email sent twice.

Give every action with an outside effect a stable reference stored in the workflow state, and use it so the destination can recognise a repeated request. Before retrying an action whose outcome is unclear, check whether it already happened. The guide to stopping an agent taking a wrong action covers how to handle uncertain writes, and what breaks first in production AI explains why silent retries are one of the first failures teams meet.

Pause properly while waiting for people

Long tasks usually wait on humans. The wrong way to handle that is to keep a process running, polling for an answer. The right way is to record exactly what the agent is waiting for, release its resources and resume only when the answer arrives or a deadline passes.

The request a person sees should stand on its own. When procurement opens an approval, they should see the proposed terms, the evidence the agent gathered and what will happen when they approve, without needing the agent's conversation history. The guide to keeping a human in the loop without friction describes how to design those checkpoints so they are quick to answer and hard to approve blindly.

Decide what happens if nobody answers. A reminder, an escalation to a named backup and eventually a clean cancellation are all better than a task that sits open indefinitely.

Recheck the world before resuming

When a task wakes up after a long pause, the facts it saved may no longer be true. The supplier's bank details might have changed, the budget might have been spent, or someone might have completed the step by hand. Acting on stale state is one of the quieter ways a long-running agent goes wrong.

Before any step with an outside effect, reload the facts that step depends on and compare them with what the agent saw earlier. If something material has changed, route the task back for a decision rather than pressing on. Approvals should be tied to the version of the proposal that was approved, so a change made after approval triggers a fresh review.

Credentials need the same care. A token issued when the task began may have expired, and the user it acts for may have lost the permission it relied on. The guide to agent authentication on behalf of a user explains why authority should be checked at the moment of action, not assumed from the start of the run.

Keep each run tied to the version that started it

Over the life of a long task the agent itself may change. A new prompt is deployed, a tool's schema is updated or the model is replaced. A run that began under one configuration and finishes under another can behave in ways nobody tested.

Record the model, prompt and tool versions with each run. For most changes, let runs already in progress finish on the version they started with and send new runs to the new version. Where that is not possible, for example when a model is withdrawn, treat the switch as a planned migration and test it against paused runs as well as new ones. The guide to what to do when a model is deprecated covers that planning.

Give operators a view of every open run

Short tasks either succeed or fail in front of someone. Long tasks can stall quietly. Operators need a list of every open run showing its current step, what it is waiting for, how long it has been waiting and whether anything unusual has happened.

From that view they should be able to pause a run, cancel it, retry a failed step or hand it to a person. Alert on runs that have waited longer than expected, runs that keep failing the same step and runs that resumed into a changed situation. The guides to monitoring an agent in production and who operates an agent after launch cover the signals and the ownership that make this work.

Because every step is saved, the same record doubles as an audit trail. The guide to auditing what an AI agent did explains what that record needs to contain when someone later asks why a decision was made.

Design for time from the start

An agent that finishes in one sitting can get away with keeping its progress in memory. An agent that works across days cannot. Saved state, repeat-safe actions, proper pauses, fresh checks before acting and version tracking are what let it survive restarts, deployments and slow humans without losing work or doing damage.

These properties are much easier to design in than to retrofit once runs are already open in production. CodeDTX's enterprise agentic AI and AI agent development work builds long-running agents on durable workflows from the first release. If you are planning an agent whose tasks outlast a single session, talk to CodeDTX about how to structure it.

Frequently asked questions

Is a long context window enough for a multi-day agent task?

No. A larger context window lets the model see more of the conversation, but it does not keep the task alive when the process restarts, stop a step running twice or notice that a saved fact has changed. Long tasks need their progress written to durable storage as structured state. The context window then only needs the parts of that state relevant to the next decision.

Do you need a workflow engine to run long agent tasks?

Not always. A simple task with a few steps can work with a well-designed state table, a queue and a scheduler that wakes runs when events arrive. As tasks grow longer, involve more people or call more systems, a durable workflow engine saves a lot of effort, because it already handles persistence, safe replay after failures, timers and waiting for external signals.

How is this different from giving an agent memory?

Memory is about what an agent knows across separate tasks, such as a user's preferences or lessons from earlier work. Durable execution is about one task surviving over time: which steps are done, what the agent is waiting for and what it must not repeat. A long-running agent needs the second even if it has none of the first, and the two should be stored and governed separately.

What should happen to a run that has been waiting too long?

Set an expected waiting time for each pause when you design the workflow. When it passes, send a reminder, then escalate to a named backup, and finally cancel the run cleanly if nobody responds. Cancellation should record why the run stopped and undo or flag any partial work, so nothing is left half finished in another system without an owner.

Contact us to build the right product

Talk to our engineers about your application, the systems it connects to, and what you want to build next.

Get in touch
Two people discussing work with a laptop