Build on GPT-6.1 Sol by default and route only your hardest tasks to GPT-6 Astra. OpenAI announced Astra on 3 September 2026 as its most capable model, then released GPT-6.1 Sol on 29 September with near-Astra performance at a fifth of the price. For most agent workloads, the cheaper model is the sensible baseline.
What OpenAI released in September
OpenAI introduced GPT-6 Astra on 3 September 2026. The launch post positions it for computer use, software engineering and professional work such as documents, spreadsheets and presentations that follow a team's own templates. Astra rolled out first to a limited set of organisations, then to ChatGPT Plus, Pro, Business and Enterprise, the OpenAI API, Microsoft Azure and AWS Bedrock.
Two behaviour changes matter for anyone running agents. Astra asks focused questions when an instruction leaves room for a consequential wrong call, and in Codex it can ask asynchronously while carrying on with work that does not depend on the answer. It is also better at keeping the original goal when a user steers it mid-task, instead of treating each new message as a fresh request.
On 22 September OpenAI added GPT-6 Sol and GPT-6 Luna to the family. On 29 September it released GPT-6.1 Sol, described in OpenAI's model documentation as near-Astra performance for complex coding, computer use and professional work at a lower cost.
How the two models compare
Both models come from the same generation, so the differences are mostly cost, endpoints and a few operational details. From OpenAI's model documentation:
- Context and output: both offer a 1,050,000-token context window, up to 922,000 input tokens and up to 128,000 output tokens, with an April 2026 knowledge cutoff.
- Price: GPT-6.1 Sol's standard input and output token prices are one-fifth of Astra's, according to the GPT-6 Astra and GPT-6.1 Sol model pages. Sol's cached input is cheaper still, which matters for agents that resend the same system prompt and tool definitions on every turn.
- Long prompts: on both models, requests above 272,000 input tokens are billed at a higher rate for the whole request. Batch and Flex processing are half the standard rate.
- Reasoning effort: both accept low, medium, high, xhigh and max. GPT-6.1 Sol defaults to medium and does not support the none or minimal settings.
- Endpoints: both work with the Responses API and Chat Completions, but neither supports Realtime, the Assistants API or fine-tuning. For GPT-6.1 Sol, tool calling needs the Responses API.
- Data residency: GPT-6.1 Sol supports US and EU data residency, with fast mode unavailable under EU residency.
When GPT-6 Astra is worth the extra cost
Astra earns its price on work where one wrong step is expensive or where the task runs for a long time without supervision. Long computer-use sessions, such as working through a web application to update records or testing a site end to end, are the clearest case: OpenAI's launch material focuses on Astra's speed and accuracy there.
The other strong case is ambiguous, high-stakes instructions. If your agent drafts client documents, touches production systems or makes decisions a reviewer will not see straight away, Astra's habit of asking before a consequential guess is worth paying for. Keep it for the steps that need that judgement, not for every call in the workflow.
When GPT-6.1 Sol is the better default
For most agent traffic - classification, retrieval, routine code changes, tool calls with clear inputs - GPT-6.1 Sol should be the starting point. The quality gap OpenAI describes is small, and the saving compounds quickly in agents that make many model calls per task.
Sol is also the simpler choice if you need EU data residency today, because OpenAI documents it for GPT-6.1 Sol. If you serve European customers, confirm the residency options for any model before you move production traffic, and read our note on data residency for AI agents.
How to switch without breaking your agent
Treat the move like any other model change, not a configuration tweak:
- Run your evaluations on both. Compare task success, tool-call errors and cost per completed task at the reasoning effort you use today. Our guide to what AI agent evals catch covers what to include.
- Route by difficulty. Send most steps to Sol and escalate to Astra only for the task types where your evaluations show Sol falling short.
- Check your endpoints. If you rely on Chat Completions for tool calling, move those calls to the Responses API before switching to Sol.
- Roll out in stages. Start with a small share of traffic and widen it as the results hold, as described in rolling out an AI agent in stages.
- Watch cost and quality in production. Track spend per task alongside success rates, using the approach in controlling what an AI agent costs to run and monitoring an AI agent in production.
Older GPT-5.6 models will eventually be retired, so plan the move now rather than under a deadline. Our note on what to do when a model is deprecated explains how to prepare.
Frequently asked questions
Is GPT-6.1 Sol as capable as GPT-6 Astra?
Not quite. OpenAI describes GPT-6.1 Sol as near-Astra performance for complex coding, computer use and professional work, and its own documentation suggests comparing the two on your tasks to judge the trade-off between quality and cost. For most agent steps the gap will not matter, but for the hardest reasoning and long computer-use sessions Astra remains the stronger option.
Do both models have the same context window?
Yes. According to OpenAI's model documentation, both GPT-6 Astra and GPT-6.1 Sol offer a 1,050,000-token context window, accept up to 922,000 input tokens and return up to 128,000 output tokens, with an April 2026 knowledge cutoff. Very long prompts are billed at a higher rate on both, so retrieving only the context a task needs still pays off.
Can we use these models from Azure or AWS?
GPT-6 Astra is available through the OpenAI API, Microsoft Azure and AWS Bedrock, as well as ChatGPT Plus, Pro, Business and Enterprise, according to OpenAI's launch announcement. If your organisation already runs workloads on one of those clouds, check the model catalogue there for GPT-6.1 Sol too, and confirm region and data residency options before you move production traffic.
How should we test a switch to GPT-6.1 Sol?
Run your existing evaluation set against both models at the reasoning effort you use today, then compare task success, tool-call errors and cost per completed task rather than headline benchmarks. Move a small share of traffic first, watch the results in production monitoring, and keep Astra available as a fallback for the task types where Sol falls short.



