OpenAI has announced pricing reductions for its GPT-5.6 model family, reducing the cost of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. The company says the changes are based on improvements across models, inference systems and agent workflows to reduce the cost and time required for AI tasks.
The lower pricing for Luna and Terra is also reflected in how usage is counted against paid subscription limits when using Codex and ChatGPT Work. GPT-5.6 Luna supports tool usage and multi-step workflows for high-volume tasks, while GPT-5.6 Terra is designed for everyday workloads.
OpenAI has also introduced Fast mode for GPT-5.6 Sol in the API, replacing Priority Processing. Fast mode provides up to 2.5× faster speeds compared with Standard processing at twice the price, without any change in intelligence. Existing API requests marked as priority will continue to work and will automatically use Fast mode.
GPT-5.6 efficiency improvements
OpenAI says GPT-5.6 improvements come from optimising models, inference systems and the agent systems that connect models with tools and context.
The improvements include:
- Better routing to improve hardware utilisation
- Optimised production software for more efficient token generation
- Improved context management to avoid repeating completed work
- Reduced time, token usage and compute required for tasks
OpenAI says GPT-5.6 Sol has helped with internal optimisation efforts by rewriting and optimising production kernels, designing and running hundreds of experiments to improve token generation, and monitoring training while intervening when problems occur.
The company says these efforts reduced the end-to-end cost of serving GPT-5.6 Sol by 20%, while experiments improved token-generation efficiency by more than 15%.
GPT-5.6 model selection and performance
OpenAI says choosing an AI model depends on factors such as the expected outcome, cost of errors, urgency and workload scale. Different tasks may require different balances of intelligence, speed, reliability and cost.
According to OpenAI, GPT-5.6 Luna delivers performance comparable to models that were frontier-class a year ago at around 6 cents per dollar per task and with nearly nine times faster execution. On the Agents’ Last Exam evaluation for professional work, OpenAI says Luna outperforms Fable 5 with an estimated cost per task that is nearly 99% lower.
The company says businesses can use evaluations to identify where additional intelligence improves results and where faster, lower-cost processing can provide the required outcome.
For example, a coding workflow could use GPT-5.6 Sol to resolve uncertainty and create a plan, then use GPT-5.6 Luna to implement defined changes, write and run tests, and evaluate results.
Pricing
Starting July 30, GPT-5.6 API pricing is:
- GPT-5.6 Terra: $2 per million input tokens and $12 per million output tokens
- GPT-5.6 Luna: $0.20 per million input tokens and $1.20 per million output tokens
- GPT-5.6 Sol: Pricing remains unchanged
ChatGPT and Codex subscription prices and quota budgets remain unchanged, while GPT-5.6 Luna and Terra usage now consumes fewer credits under paid subscription plans.
Availability
GPT-5.6 Terra and Luna are available through ChatGPT Work, Codex and the OpenAI API. In ChatGPT Work and Codex, Free and Go users can access Terra, while Plus, Pro, Business and Enterprise users can choose Terra and Luna.
The updated pricing will begin rolling out to AWS. Fast mode for GPT-5.6 Sol is available through the API and aligns with /fast in Codex.