Openai has launched Ultrafast, a new service tier running GPT-5.6 Sol up to 14 times faster than Standard processing in the OpenAI API, powered by Cerebras to generate up to 750 output tokens per second, according to the company. The release brings advanced intelligence to time-sensitive workflows where every second matters, eliminating the traditional trade-off between speed and model capability.
Cerebras Partnership Powers Ultra-Low-Latency Inference
The new service tier marks the next step in OpenAI’s partnership with Cerebras to bring ultra-low-latency inference to the platform. By leveraging Cerebras hardware, GPT-5.6 Sol on Ultrafast mode delivers up to 750 output tokens per second, according to OpenAI. Until now, achieving real-time speed usually meant choosing a smaller or more specialized model. Ultrafast allows businesses to utilize frontier intelligence at high speeds, enabling more useful work per second across demanding operational environments.
Early Business Workflows and Customer Testing
OpenAI is testing GPT-5.6 Sol on Ultrafast mode with an initial group of companies across coding, commerce, financial research, support, and other interactive applications. Initial testing partners include Jane Street, Podium, Basis, and Rogo, according to company disclosures. During the preview period, OpenAI is working with this group to understand where order-of-magnitude speed changes create the most value and how products change when models keep pace with human users. Businesses requiring frontier intelligence at high speeds can sign up to get notified when access expands.
Internal Use Cases in Incident Response and Research
Inside OpenAI, internal developer teams are testing GPT-5.6 Sol on Ultrafast mode to evaluate real-time responsiveness. In incident response, when critical systems fail, engineers use the service to read logs, analyze traces, synthesize conversations, and prepare or validate fixes while outages unfold. For research, teams use Ultrafast to search knowledge sources, query data, and synthesize information across connected tools, turning overnight batch runs into interactive working sessions during the regular workday.
Pro Tip: When evaluating real-time AI for incident response or financial security, focus on workflows where reducing the delay between observing a signal and testing a hypothesis improves operational outcomes while keeping humans responsible for final deployment.
Frequently Asked Questions
What is Ultrafast mode?
Ultrafast is a new OpenAI API service tier running GPT-5.6 Sol up to 14 times faster than Standard processing, generating up to 750 output tokens per second through a partnership with Cerebras.
Who is currently using Ultrafast?
An initial group of preview customers—including Jane Street, Podium, Basis, and Rogo—is testing the service across coding, commerce, financial research, and support workflows, alongside internal OpenAI developer teams.
What are the primary use cases for the service?
Primary use cases include incident response, financial research and security, customer support and voice applications, commerce checkouts, and live research experimentation.
How can businesses access Ultrafast?
Businesses requiring frontier intelligence at high speeds can sign up on OpenAI’s platform to receive notifications when access expands beyond the initial preview period.
Join the Discussion: How could real-time frontier intelligence reshape your organization’s workflow? Share your thoughts in the comments below, explore our related AI architecture analyses, or subscribe to our newsletter for the latest infrastructure updates.
Keep reading