AWS has released an agent skill, aws-ai-ml, that lets coding agents such as Claude Code, Codex and Kiro benchmark Amazon SageMaker endpoints, recommend instance types and write SageMaker Python SDK v3 code. For teams hosting their own models, it turns inference sizing into a conversation, but the load tests and endpoints it creates still cost real money.
What AWS released
The skill was announced by AWS on 5 October 2026 in the post New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent. It is part of the Agent Toolkit for AWS, which bundles the AWS MCP Server, skills, plugins and rule files for coding agents.
There are two ways to get it:
- Any coding agent. Run
aws configure agent-toolkit, thennpx skills add aws/agent-toolkit-for-aws/skills/aws-ai-ml. AWS names Kiro, Claude Code and Codex, plus any MCP-compatible agent. - SageMaker Studio. Create a JupyterLab space with the pre-configured image that already includes the skill, then sign in to your coding agent from the terminal.
The prerequisites are AWS CLI 2.35 or later, the uv package manager, and AWS credentials that can call the SageMaker APIs for creating endpoints and running benchmark and recommendation jobs.
What actually changed
The underlying capability is not new. In April 2026, AWS added optimised generative AI inference recommendations to SageMaker AI: you describe a workload, choose to optimise for cost, latency or throughput, and SageMaker benchmarks candidate configurations on real GPUs and returns ranked options. Until now, that meant working through the APIs and sample notebooks.
The new skill puts the same work behind a coding agent. According to AWS, the agent can:
- Benchmark an existing endpoint, reporting throughput, latency percentiles, time to first token, inter-token latency and concurrency.
- Recommend an instance type for a fine-tuned model in Amazon S3, a SageMaker JumpStart foundation model, or a Hugging Face Hub model, including gated ones.
- Compare two benchmark runs and show the difference on each metric.
- Generate SageMaker Python SDK v3 code that you can read, change and run yourself.
AWS also describes guard rails in how the agent behaves: it asks for missing details instead of guessing, it confirms that an endpoint is safe to load-test before sending traffic to it, and it says when a request, such as deploying a model, is outside what the skill does.
What it costs
The Agent Toolkit has no charge of its own; AWS says you pay for the resources your agent provisions or uses. The April announcement says generating recommendations has no additional cost, but the optimisation jobs and benchmarking endpoints are billed at normal compute rates. AWS's clean-up steps tell you to delete endpoints, stop the JupyterLab space and remove benchmark output from your default S3 bucket to avoid ongoing charges.
In CodeDTX's view, that is the part to plan for. An agent that can start GPU benchmarks on request makes it easy to run more of them than anyone meant to. Our guide to controlling what an AI agent costs to run applies directly: set a budget alert on the account the agent uses, and tag what it creates so you can find and delete it.
What it means for teams building AI products
This skill matters mostly to teams that already run, or are weighing, an AI agent on a model they host themselves. Picking an instance type for a self-hosted model usually means writing and running your own load tests; this makes a first pass something an engineer can ask for in plain language. Our practical advice:
- Give the agent its own IAM role. The skill needs permission to create endpoints and run jobs. Scope that role to a sandbox account or a tagged set of resources, never production, as in how to stop agents reaching restricted data.
- Never point it at a live endpoint without a second check. The agent asks before load-testing, but a "yes" typed by a busy engineer is not a change control. Benchmark a copy of the endpoint instead.
- Keep the generated code. Because the output is SDK code, commit it alongside your deployment configuration, so the instance choice is reviewable and repeatable rather than buried in a chat history.
- Treat skills as supply chain. The skill is installed from a public repository and runs inside your agent with your credentials. Review it like any other dependency, as our note on AI agent skills security and governance explains.
When not to adopt it
Skip it if your models run on Amazon Bedrock or a hosted API rather than on SageMaker endpoints, because there is nothing for it to size. Hold back too if you have no separate AWS account or budget controls for experiments: the risk is not a wrong recommendation, it is GPU endpoints left running. And if your team already has a tested load-testing harness, use the skill to compare against your own numbers before trusting its ranking.
Frequently asked questions
What is the Amazon SageMaker aws-ai-ml agent skill?
It is an agent skill from AWS, announced on 5 October 2026, that teaches coding agents to work with SageMaker AI inference. Once installed, an agent can benchmark an existing endpoint, recommend instance types for a model, compare benchmark runs and generate SageMaker Python SDK v3 code. It is part of the Agent Toolkit for AWS and works with Kiro, Claude Code, Codex and MCP-compatible agents.
How do you install the SageMaker inference skill?
Locally, run aws configure agent-toolkit and then npx skills add aws/agent-toolkit-for-aws/skills/aws-ai-ml, with AWS CLI 2.35 or later and uv installed. In SageMaker Studio, create a JupyterLab space with the pre-configured image that includes the skill and sign in to your coding agent from the terminal. Either way, your AWS credentials need permission to call the SageMaker APIs.
Does the SageMaker agent skill cost anything?
The Agent Toolkit for AWS has no charge of its own, and AWS says generating inference recommendations has no additional cost. You do pay standard compute prices for the optimisation jobs and benchmarking endpoints the agent starts. Delete endpoints, stop Studio spaces and clear benchmark output from S3 when you finish, and put a budget alert on the account the agent uses.
Can the skill deploy a model to production?
No. AWS says that if you ask for something outside its scope, such as deploying a model, the agent explains what it cannot do. The skill benchmarks, recommends and writes SDK code that you run yourself. That keeps the deployment step with your team, which is where it belongs, along with your normal review and change control.



