OpenAI rolls out GPT-6.1 Sol with improved performance and $2 per million input pricing

OpenAI has introduced GPT-6.1 Sol as an upgrade to GPT-6 Sol. The company says the new model nearly matches GPT-6 Astra’s intelligence on agentic coding, computer use, and professional work.

GPT-6.1 Sol: Coding, professional work and computer use

The model improves over GPT-6 Sol across writing and debugging code, understanding documents, and executing multi-step business workflows. OpenAI evaluated it across software engineering, professional document understanding, business automation, and computer-use tasks, with the following results:

  • DeepSWE v1.1: This evaluation covers complex software-engineering tasks in real codebases. The model matches GPT-6 Astra at roughly one-fifth the cost and scores 6.4 percentage points higher than GPT-6 Sol at a lower reasoning effort and cost.

  • GDP.pdf: This evaluation measures how accurately models answer professional questions using complex PDF documents containing tables, charts, diagrams, and fine-print details. The model scores higher than Opus 5.5 with fallbacks at less than half the cost per task across the tested reasoning settings. It also approaches GPT-6 Astra’s state-of-the-art performance at roughly one-fifth the cost per task.

  • AutomationBench: This evaluation measures whether agents correctly complete multi-step business workflows. The model scores 2.2 percentage points above Opus 5.5 at medium reasoning effort at roughly one-third the cost. It also scores 4.8 percentage points higher than GPT-6 Sol at the same setting.

  • OSWorld 2.0: On the offline set for demanding computer-use workflows, the model outperforms GPT-6 Sol by seven percentage points at maximum reasoning effort at less than half the cost. It comes within 2.1 percentage points of GPT-6 Astra’s score at the same reasoning setting at roughly one-seventh the cost per task.

Scientific research and factuality

On Terminal-Bench Science 0.1, which evaluates scientific workflows including data analysis, simulation, and theorem proving, the model more than doubles GPT-6 Sol’s score at maximum reasoning effort while costing less than half as much per task.

At maximum reasoning effort, the average cost per task is:

  • GPT-6.1 Sol: $5.47
  • Opus 5.5: $23.21
  • GPT-6 Astra: $23.80

The cost per task is therefore more than 75% lower than both Opus 5.5 and GPT-6 Astra in this evaluation. GPT-6 Astra achieves the highest score among the tested models at 68.1%, and OpenAI says it should be used for the most difficult scientific research tasks.

For factuality, the share of responses containing a factual error falls from 11.4% with GPT-6 Sol to 7.7% at low reasoning effort, a reduction of approximately 32%. Across the tested reasoning settings, the error rate remains within 1.9 percentage points of GPT-6 Astra’s at less than one-fifth the cost per task.

The factuality evaluation measures the share of answers containing at least one factual error on de-identified conversations where users had flagged an earlier model’s error. OpenAI notes that these deliberately difficult prompts are not representative of typical usage.

GPT-6.1 Sol: Safety and alignment

The model shows improvements over GPT-6 Sol in OpenAI’s alignment evaluations, bringing it closer to GPT-6 Astra. OpenAI says it is more transparent about its limitations and more reliable at respecting user intent and safety constraints.

In challenging evaluations, it shows lower failure rates than GPT-6 Sol in the following areas:

  • Transparency about broken search tools
  • Respecting explicit restrictions
  • Avoiding unauthorized outcomes during agentic tasks

OpenAI observed no attempts to bypass an automated safety reviewer, matching GPT-6 Astra and GPT-6 Sol. The company notes that these evaluations deliberately test challenging situations and do not measure failure rates in typical use. Full details are available in the GPT-6.1 Sol system card addendum.

Pricing and availability

GPT-6.1 Sol is available starting today to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. It is not yet available in Chat.

Developers can access the model through the OpenAI API as gpt-6.1-sol. Its standard API prices are:

  • Input: $2 per million tokens
  • Cached input: $0.10 per million tokens
  • Output: $10 per million tokens

The cached input price is 95% lower than standard input pricing and 50% lower than GPT-6 Sol’s cached input pricing. OpenAI says this gives developers more room to build and run agents that reuse context across requests.

In the coming days, OpenAI will also offer GPT-6.1 Sol Ultrafast, with up to 8x faster token generation compared with its standard speed in Codex.


Related Post