Skip to main content

Should you run an AI agent on a model you host yourself?

Open-weight models are now capable enough to run many agents. Hosting one yourself buys data control and stable versions, but you take on serving, security, evaluation and upgrades. Here is how to decide.

An engineer types on a laptop in a large server hall while a colleague faces monitors showing a glowing blue AI brain.

Run an AI agent on a model you host yourself when data rules, predictable high volume or the need to control model changes outweigh the work. Hosting shifts effort from paying a provider to running GPUs, serving, security, evaluation and upgrades. Many teams end up with both: hosted models for sensitive or steady work, provider APIs for the rest.

Why this is now a real choice

Until recently the answer was usually simple. The models you could download and run were noticeably weaker than the ones offered through provider APIs, so any agent that had to plan, call tools and recover from mistakes ran on a hosted frontier model.

That gap has narrowed. Open-weight models, where the trained weights are published and can be run on your own hardware or in your own cloud account, now handle tool calling, structured output and multi-step reasoning well enough for a large share of business agents. Serving software has matured too, so running a model behind an internal endpoint no longer needs a research team.

That turns a technical default into a business decision. The question is no longer whether a self-hosted model can do the job at all, but whether the benefits of running it yourself are worth what it costs your team to operate.

What you gain by hosting the model

The first gain is data control. Prompts, documents, tool results and outputs stay inside infrastructure you govern. There is no third-party retention policy to review and no cross-border transfer to justify. For agents that read contracts, patient records, source code or customer data under strict rules, that can decide the matter on its own. The guide to stopping an agent reaching data it should not see covers the controls that still apply either way.

The second is control over change. A provider can update or withdraw a model on its own timetable. A model you host changes only when you decide to change it, which makes behaviour easier to keep stable and approvals easier to defend.

The third is cost shape. Provider APIs charge per token, so spend rises with usage. Self-hosting is mostly a fixed cost for capacity. For steady, heavy workloads that can work out cheaper. For spiky or light workloads it usually does not, because you pay for capacity whether or not it is used.

Some teams also gain lower latency, because the model sits next to their systems, and the ability to run where there is no reliable internet connection.

What you take on

Hosting a model means running a production service with unusual hardware. You need GPU capacity, sized for peak demand rather than average demand, plus a serving layer that batches requests, handles queues and scales. Someone has to patch it, secure the endpoint, rotate credentials and respond when it slows down at the worst moment.

You also inherit work a provider normally does quietly. That includes safety filtering, abuse monitoring, capacity planning and choosing when to move to a newer model. Each upgrade needs the same testing as any other release.

This is ongoing operational work, not a one-off setup, and it needs named owners. The guides to who operates an agent after launch and monitoring an agent in production describe the responsibilities involved, and a self-hosted model adds the model service itself to that list.

Check the model can do the agent's job

General benchmark scores say little about whether a model will run your agent well. What matters is how it performs on your tasks: choosing the right tool, producing arguments that match the schema, following your policies, handling your documents and recovering when a tool returns an error.

Build an evaluation set from real examples of the agent's work before you commit, and run the candidate open-weight model and your current hosted model against it side by side. Look closely at the failures, not just the pass rate. A model that is slightly less accurate but fails safely may be the better choice. The guide to what evals for AI agents actually catch explains how to build that set.

Check the quantised version you will actually serve, not only the full-size model. Compressing a model to fit cheaper hardware can change its behaviour in ways that only show up on harder tasks.

Read the licence and know where the model came from

Open-weight does not always mean free to use however you like. Licences vary: some allow commercial use without conditions, others restrict certain uses, require attribution or apply limits to large deployments. Your legal team should read the licence for the exact model and version you plan to run.

Treat the model file like any other third-party component. Download it from the publisher's official source, record its version and checksum, and review what is known about how it was trained. The same supply-chain thinking described in the guide to governing AI agent skills applies to the weights themselves. For regulated work, the guide to compliance with non-deterministic systems covers the evidence reviewers will expect.

Design so the model can be swapped

Whichever route you choose, avoid wiring the agent tightly to one model. Put a thin internal interface between the agent and the model, so prompts, tool definitions and output validation do not depend on one provider's quirks. Then moving from a hosted API to a self-hosted model, or back, becomes a configuration change backed by evals rather than a rebuild.

That same layer lets you route by task. A hypothetical claims-handling agent might send anything containing personal data to a self-hosted model inside the company's network, while general drafting and summarising go to a provider API. Routing also helps with cost, as the guide to controlling what an agent costs to run explains, and with model changes, covered in what happens when a model is deprecated.

A practical way to decide

Work through a few questions in order. Does a regulation, contract or internal policy stop this data leaving your infrastructure? If so, self-hosting or a private deployment is likely required, and the rest is about doing it well. If not, is the workload steady and heavy enough that fixed capacity would be well used? Do you have, or can you build, a team to run GPU services properly? And does an open-weight model pass your evals for this agent?

If the answers are mostly no, start with a provider API and keep the swap layer in place. If they are mostly yes, pilot the self-hosted model on one agent, compare it with the hosted version in real use and expand from there. The guide to rolling out an agent in stages describes how to run that comparison safely.

Choose the hosting model deliberately

Self-hosting is not a badge of seriousness, and provider APIs are not a shortcut to avoid. Each is a trade between control and operational effort. The right choice depends on your data, your workload and your team, and it can differ between agents in the same company.

CodeDTX's enterprise AI integration and AI data boundaries work helps teams make that call with evidence, and build agents that can move between hosted and self-hosted models without a rewrite. If you are weighing where your agent's model should run, talk to CodeDTX.

Frequently asked questions

Is a self-hosted model always cheaper than a provider API?

No. Self-hosting replaces a per-token bill with the fixed cost of GPU capacity and the people who run it. That tends to pay off only when usage is steady and heavy enough to keep the hardware busy. For light or unpredictable workloads, paying per request is usually cheaper, because idle capacity still costs money and the operational work does not shrink with usage.

Are open-weight models good enough for AI agents?

For many business agents, yes. Current open-weight models handle tool calling, structured output and multi-step tasks far better than earlier releases. Whether one is good enough for your agent depends on your tasks, not general rankings. Run your own evaluation set against the exact version you plan to serve, including any compressed variant, and compare its failures with your current model before deciding.

Does self-hosting remove the need for data governance?

No. Keeping the model inside your infrastructure removes the risk of data going to a third party, but the agent can still reach data it should not, leak information between users or take actions beyond its authority. Access controls, data boundaries, audit trails and approval steps are needed whether the model runs on your servers or behind a provider API.

Can we use a provider API and a self-hosted model together?

Yes, and many teams do. Put a common internal interface in front of both, then route each request by its data sensitivity, cost or quality needs. Sensitive work can stay on the self-hosted model while general tasks use a provider. Keep separate evals for each route, and log which model handled every request so you can explain and audit the agent's behaviour later.

Contact us to build the right product

Talk to our engineers about your application, the systems it connects to, and what you want to build next.

Get in touch
Two people discussing work with a laptop