OpenAI has previewed Ultrafast, a new service tier that runs GPT-5.6 Sol at up to 14× the speed of Standard processing, launching first through the OpenAI API.
GPT-5.6 Sol Ultrafast
Powered by Cerebras for low-latency inference, Ultrafast can generate up to 750 output tokens per second. OpenAI says improvements to GPT-5.6 have increased efficiency across its stack, while Ultrafast is being tested for workflows that require faster responses from the model.
Until now, applications requiring real-time AI responses often relied on smaller or more specialized models. OpenAI is testing Ultrafast across several types of workloads:
- Incident response and reliability: Analyze application logs, recent code changes and engineer reports to identify a likely cause and help prepare a fix while an outage is still in progress.
- Financial research and security: Analyze market signals, assess transactions and identify suspicious activity while conditions are changing.
- Customer support and voice: Resolve complex customer issues in real time when finding an answer requires multiple steps or systems.
- Commerce: Answer product questions, check inventory, personalize recommendations and resolve checkout issues while customers are still shopping.
- Live research and experimentation: Run experiments, review results, adjust the approach and start another iteration during the workday instead of waiting for an overnight run.
Early customer testing
OpenAI is testing GPT-5.6 Sol with Ultrafast with an initial group of companies across coding, commerce, financial research, support and other interactive applications. These business workflows are being tested in real production environments to understand where the increase in speed makes the biggest difference and how products change when the model can respond more quickly.
OpenAI says these findings will help guide deployment as capacity grows.
How OpenAI is using Ultrafast
OpenAI is also testing Ultrafast internally for incident response and research. When an alert is triggered, engineers can use GPT-5.6 Sol to:
- Read application logs and analyze traces
- Synthesize relevant conversations
- Identify the next checks
- Help prepare or validate a fix
This reduces the time between observing a signal, testing a hypothesis and choosing the next action, while engineers remain responsible for judgment and deployment.
For research, OpenAI uses Ultrafast to search knowledge sources, query data, and gather, organize and summarize information across connected tools. The faster processing allows researchers to shorten the gap between running experiments and reviewing their results, supporting multiple iterations during the workday.
Availability
GPT-5.6 Sol with Ultrafast mode is currently available as a limited preview to a select group of customers through the OpenAI API. OpenAI plans to expand access as capacity grows and is accepting sign-ups for updates.