GLM 5.3, Z.ai's open-weights model for coding and long-running agent work, is now a managed model on Amazon Bedrock. For teams building AI agents on AWS, it means trying a strong open model through the APIs and account controls they already use, without hosting it, as long as their account is eligible.
What was released
GLM 5.3 on Amazon Bedrock was announced by Amazon Web Services on 5 October 2026 in the post Introducing GLM 5.3 on Amazon Bedrock. The model comes from Z.ai (Zhipu AI), which published the open weights in its GLM-5.3 model card on Hugging Face in August. What is new this week is the managed version: AWS runs the model, and you call it like any other Bedrock model.
AWS describes GLM 5.3 as a 753B-parameter mixture-of-experts model tuned for coding and long-horizon agentic tasks, with a focus on security work. AWS states that "Access to GLM 5.3 on Bedrock is available to eligible enterprise customers", so check your account before planning around it.
What actually changed
Compared with GLM 5.2, Z.ai's model card reports a large gain on its in-house Z.ai Code Bench and stronger results on public coding and agent benchmarks. Those are Z.ai's own figures, so treat them as a reason to test, not as a result for your workload.
On the Bedrock side, the announcement lists these details:
- Two ways to call it. OpenAI-compatible Responses and Chat Completions APIs, or Bedrock's own Invoke and Converse APIs. Authentication uses Bedrock API keys or short-lived AWS credentials.
- Cross-Region inference only. The model is offered through a US profile,
us.zai.glm-5.3, and a Global profile,global.zai.glm-5.3. - Service tiers. Flex for lower cost, Priority for lower latency, and Standard in between.
- Prompt caching. Both implicit (automatic) and explicit cache controls, which AWS points out matters for agents that resend large system prompts or repository context on every turn.
- Tool use and multi-step reasoning, shown in the post with Strix, an open-source penetration testing agent, run against the deliberately vulnerable OWASP Juice Shop app.
Z.ai's model card also refers to a 1M-token context window and a reasoning_effort setting with low, high and max levels. The AWS post does not state the context window or per-token prices on Bedrock, so confirm both in the Bedrock console before you size a workload.
What it means for teams building AI products and agents
In CodeDTX's view, the main change is not the model but where it runs. Until now, using GLM 5.3 meant hosting a very large model yourself, which is the trade-off we cover in should you run an AI agent on a model you host yourself. On Bedrock it becomes another model ID behind the same IAM policies, logging and billing as the rest of your AWS estate.
- Low-cost trials if you already use the OpenAI SDK. Because Bedrock exposes OpenAI-compatible endpoints for GLM 5.3, a coding agent written against the Chat Completions or Responses API can often be pointed at it by changing the base URL, key and model name. That makes a side-by-side evaluation cheap to set up. Use the method in what AI agent evals catch rather than vendor benchmarks.
- Cost control for agent loops. Prompt caching and the Flex tier are the two levers to try first for long coding sessions that resend the same context. Measure cache hit rates on your own traces, as described in our guide to controlling what an AI agent costs to run.
- Governance stays in one place. A managed model inherits your existing AWS access controls and audit trail. That is simpler than approving a new external API vendor, but it does not replace your own records of what the agent did; see what an AI agent audit trail contains.
- Security agents need extra guardrails. AWS and Z.ai both highlight cybersecurity capability. A model that is good at finding vulnerabilities should only point at systems you own, in a sandbox, with a human approving anything that touches production.
When not to switch
Stay on your current model if your account is not eligible, if you need data processed in a single named Region (the model is offered only through US and Global cross-Region profiles), or if your agent depends on features you have not tested on GLM 5.3. Check Z.ai's licence terms with your legal team before relying on the open weights as a fallback. And if your current model already passes your evals at an acceptable cost, a new model is a project, not a free upgrade: plan the change as in rolling out an AI agent in stages.
Frequently asked questions
What is GLM 5.3?
GLM 5.3 is an open-weights mixture-of-experts model from Z.ai, also known as Zhipu AI, tuned for coding and long-running agent tasks, with a stated focus on security work. Z.ai published the weights on Hugging Face in August 2026, and Amazon Web Services made a managed version available on Amazon Bedrock on 5 October 2026 for eligible enterprise customers.
How do I call GLM 5.3 on Amazon Bedrock?
You can use the OpenAI-compatible Responses or Chat Completions APIs, or Bedrock's Invoke and Converse APIs, with the model ID global.zai.glm-5.3 or the US profile us.zai.glm-5.3. Authentication uses Bedrock API keys or short-lived AWS credentials. The AWS console playground is also available for quick tests before you write any code or change an existing agent.
Does GLM 5.3 run in a single AWS Region?
The announcement lists two cross-Region inference profiles, one for the US and one Global, so requests can be served from more than one Region. If your data must stay in a single named Region for contractual or regulatory reasons, confirm with AWS where requests are processed before you send production data, and keep your current model until that is settled.
Should we move our coding agent to GLM 5.3?
Only after a side-by-side test on your own tasks. Run your current model and GLM 5.3 on the same set of real tickets or repository changes, compare success rates, cost per task and latency on the tier you would use, and check tool-calling behaviour. Switch where it matches or beats what you have, and keep the old model available for rollback.
How is this different from hosting GLM 5.3 ourselves?
On Bedrock, AWS runs the hardware and you pay per token, with AWS access controls, service tiers and prompt caching included. Hosting the open weights yourself gives full control over where the data goes and which version runs, but you buy or rent the GPUs, keep the serving stack up to date and handle scaling. A practical order is to test on Bedrock first and decide on hosting later.



