OpenAI is testing a new “Ultrafast” mode that could make its most capable AI feel far less like a system waiting for an answer and more like software that can keep up with a live conversation. The company says the early service can run GPT-5.6 Sol at up to 14 times the speed of its standard processing, though for now it is an API-only limited preview for select business customers – not a feature ordinary ChatGPT users can switch on today.
OpenAI’s next race is latency
For much of the generative AI boom, the headline competition has been about raw capability: which model writes better code, reasons through harder questions, follows instructions more reliably, or makes fewer mistakes. OpenAI’s Ultrafast preview points to a different, increasingly important contest: how quickly a frontier model can deliver that intelligence once a person hits send.
That distinction matters more than it may sound. A model can be impressive yet still feel awkward in a real working session if every request becomes a pause – particularly when it needs to read a long technical log, draft a multi-step response, or iterate on code. When responses arrive with very low delay, the interaction changes. The user is less likely to batch up work and wait; instead, they can ask, correct, test, and continue in the same flow.
OpenAI says Ultrafast runs its GPT-5.6 Sol model at up to 750 output tokens per second, a claimed maximum that is up to 14 times faster than its standard processing. The company frames the offering as a new service tier for situations where “every second matters,” rather than a smaller, stripped-down model built purely for speed.
That framing is important. Fast AI has often meant accepting a less capable model for everyday jobs, while larger and more capable models were reserved for tasks where users could tolerate a delay. Ultrafast is OpenAI’s attempt to narrow that trade-off: retain frontier-level reasoning and quality, while making the response feel immediate enough for time-sensitive work.
Not a ChatGPT toggle – yet
Despite the obvious consumer appeal, OpenAI is not currently positioning Ultrafast as a broad ChatGPT setting. The initial rollout is through the OpenAI API, with access limited to a small group of customers as the company studies where the performance difference actually creates value. OpenAI says access will expand as capacity grows, but it has not announced consumer availability, a public launch date, or pricing.
That makes the title “OpenAI testing Ultrafast mode for ChatGPT” slightly more complicated than it first appears. It is best understood as technology that could eventually shape ChatGPT experiences, but is presently being trialed in business and developer products built on OpenAI’s platform.
The company’s early focus is practical rather than flashy: coding tools, customer support, commerce, financial research, incident response, and other interactive applications. These are settings where an answer that arrives a few seconds earlier can alter the workflow itself, not merely make it feel nicer.
Take incident response. When a production system breaks, engineers may need to sift through logs, traces, recent code changes, internal messages, and monitoring alerts while the issue is still unfolding. OpenAI says its own teams are using the mode to accelerate that loop: gathering evidence, proposing next checks, and helping prepare or validate a fix. The final decision and deployment still sit with humans, but a model that can keep pace with the investigation may reduce the idle gaps between each step.
Customer support offers another clear example. A voice or chat agent that pauses for several seconds while looking up policy details, stock availability, account history, or troubleshooting instructions makes the experience feel automated in the least flattering way. A high-speed model could make a more sophisticated assistant viable during a live conversation, rather than forcing companies to choose between a quick but simplistic bot and a better but slower one.
Why 750 tokens matter
“Tokens per second” is a technical metric, but its real-world meaning is reasonably straightforward. Tokens are the small chunks of text a model uses to generate language. Faster token output means the answer appears much more quickly and long responses can be completed with less waiting.
At OpenAI’s stated peak of 750 output tokens per second, a model could theoretically stream a substantial answer almost instantly. In practice, the experience will still depend on factors such as prompt size, tool calls, network latency, model reasoning, output length, traffic levels, and the exact application design. “Up to” is doing real work in the announcement; it is a peak claim, not a promise that every request will hit the same speed.
Still, the central idea is compelling. The biggest gains may not show up in a one-off prompt asking for an email draft. They are more likely to appear in repeated loops: a developer fixing a bug, an analyst testing a hypothesis, a researcher refining a search, or a support representative solving a messy problem across several systems.
OpenAI describes one research workflow in which teams might previously launch a batch of experiments overnight and check results the next morning. With much lower latency, the company argues that the same work can become an interactive cycle during the day: run an experiment, inspect the result, revise the approach, and run another.
That is a more consequential promise than “AI answers faster.” It suggests a future where the boundary between thinking, testing, and acting becomes thinner. The useful model is not simply the one that gives the best answer eventually, but the one that can join the user’s working rhythm without constantly interrupting it.
Cerebras is central
The hardware partner behind the preview is Cerebras, a company known for specialized AI systems designed around wafer-scale computing. OpenAI says its Ultrafast mode is powered by Cerebras as part of a partnership focused on ultra-low-latency inference – the process of running a trained model to generate an answer.
This is a meaningful detail because it shows that the next stage of AI competition is not only about model training. Training the next giant model remains expensive and strategically important, but serving that model efficiently at scale is equally critical. A highly capable model that is too slow, too costly, or too constrained to deploy broadly can struggle to become a genuinely useful product.
Cerebras says the implementation preserves GPT-5.6 Sol’s intelligence while delivering the speed improvement, although such claims should be judged over time through independent testing and real-world customer use. For now, OpenAI is explicitly treating Ultrafast as a preview, which is an acknowledgement that performance, capacity, reliability, and product design all need to be proven beyond a launch-day benchmark.
A shift in what AI can do
It is tempting to see Ultrafast as a simple infrastructure upgrade, the kind of technical change most people never notice. But speed can alter what is economically and ergonomically possible.
For a shopping assistant, latency may determine whether a customer stays engaged long enough to get an answer about fit, availability, delivery, or returns. For fraud detection, the value of a signal may fall quickly as a transaction progresses. For a developer tool, a slow assistant encourages context switching; a fast one may become part of the actual editing and debugging loop. OpenAI identifies commerce, financial research, security, coding, support, and live experimentation as the areas it is watching most closely during the trial.
The company is also testing the technology with an initial group of customers including Jane Street, Podium, Basis, and Rogo. John Crepezzi of Jane Street said the speed made it practical for developers to work “in a more focused and productive way” alongside models, which captures the broader thesis behind the launch: reducing delay can change behavior, not just benchmark charts.
For ChatGPT users, the immediate takeaway is more modest. There is no public Ultrafast button to enable, and no guarantee the exact tier will arrive in the consumer product. But OpenAI’s decision to test it offers a strong hint about where the company sees the next product advantage: not only smarter answers, but AI that can respond quickly enough to feel present in the moment.
If that idea holds up outside the limited preview, the real breakthrough may not be a chatbot that types faster. It may be a shift toward AI systems that are capable enough for difficult work and responsive enough that people actually want to keep them open while they do it.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
