GadgetBond

  • Latest
  • How-to
  • Tech
    • AI
    • Amazon
    • Apple
    • CES
    • Computing
    • Creators
    • Google
    • Meta
    • Microsoft
    • Mobile
    • Samsung
    • Security
    • Xbox
  • Transportation
    • Audi
    • BMW
    • Cadillac
    • E-Bike
    • Ferrari
    • Ford
    • Honda Prelude
    • Lamborghini
    • McLaren
    • Mercedes
    • Porsche
    • Rivian
    • Tesla
  • Culture
    • Apple TV
    • Disney
    • Gaming
    • Hulu
    • Marvel
    • HBO Max
    • Netflix
    • Paramount
    • SHOWTIME
    • Star Wars
    • Streaming
Add GadgetBond as a preferred source to see more of our stories on Google.
Font ResizerAa
GadgetBondGadgetBond
  • Latest
  • Tech
  • AI
  • Deals
  • How-to
  • Apps
  • Mobile
  • Gaming
  • Streaming
  • Transportation
Search
  • Latest
  • Deals
  • How-to
  • Tech
    • Amazon
    • Apple
    • CES
    • Computing
    • Creators
    • Google
    • Meta
    • Microsoft
    • Mobile
    • Samsung
    • Security
    • Xbox
  • AI
    • Anthropic
    • ChatGPT
    • ChatGPT Atlas
    • Gemini AI (formerly Bard)
    • Google DeepMind
    • Grok AI
    • Microsoft Copilot
    • OpenAI
    • Perplexity
    • xAI
  • Transportation
    • Audi
    • BMW
    • Cadillac
    • E-Bike
    • Ferrari
    • Ford
    • Honda Prelude
    • Lamborghini
    • McLaren W1
    • Mercedes
    • Porsche
    • Rivian
    • Tesla
  • Culture
    • Apple TV
    • Disney
    • Gaming
    • Hulu
    • Marvel
    • HBO Max
    • Netflix
    • Paramount
    • SHOWTIME
    • Star Wars
    • Streaming
Follow US
AIOpenAITech

OpenAI tests Ultrafast GPT-5.6 Sol with Cerebras

The company says the early service can run GPT-5.6 Sol at up to 14 times the speed of its standard processing, though for now it is an API-only limited preview for select business customers.

By
Shubham Sawarkar
Shubham Sawarkar's avatar
ByShubham Sawarkar
Editor-in-Chief
I’m a tech enthusiast who loves exploring gadgets, trends, and innovations. With certifications in CISCO Routing & Switching and Windows Server Administration, I bring a sharp...
Follow:
- Editor-in-Chief
Aug 14, 2026, 9:00 AM EDT
Share
We may get a commission from retail offers. Learn more
Split-screen comparison of GPT-5.6 Sol Ultrafast and Standard modes building a 3D warehouse simulator, with Ultrafast marked “Task Complete” and showing more verified build events and tests passed.
Image: OpenAI
SHARE

OpenAI is testing a new “Ultrafast” mode that could make its most capable AI feel far less like a system waiting for an answer and more like software that can keep up with a live conversation. The company says the early service can run GPT-5.6 Sol at up to 14 times the speed of its standard processing, though for now it is an API-only limited preview for select business customers – not a feature ordinary ChatGPT users can switch on today.

OpenAI’s next race is latency

For much of the generative AI boom, the headline competition has been about raw capability: which model writes better code, reasons through harder questions, follows instructions more reliably, or makes fewer mistakes. OpenAI’s Ultrafast preview points to a different, increasingly important contest: how quickly a frontier model can deliver that intelligence once a person hits send.

That distinction matters more than it may sound. A model can be impressive yet still feel awkward in a real working session if every request becomes a pause – particularly when it needs to read a long technical log, draft a multi-step response, or iterate on code. When responses arrive with very low delay, the interaction changes. The user is less likely to batch up work and wait; instead, they can ask, correct, test, and continue in the same flow.

OpenAI says Ultrafast runs its GPT-5.6 Sol model at up to 750 output tokens per second, a claimed maximum that is up to 14 times faster than its standard processing. The company frames the offering as a new service tier for situations where “every second matters,” rather than a smaller, stripped-down model built purely for speed.

That framing is important. Fast AI has often meant accepting a less capable model for everyday jobs, while larger and more capable models were reserved for tasks where users could tolerate a delay. Ultrafast is OpenAI’s attempt to narrow that trade-off: retain frontier-level reasoning and quality, while making the response feel immediate enough for time-sensitive work.

Not a ChatGPT toggle – yet

Despite the obvious consumer appeal, OpenAI is not currently positioning Ultrafast as a broad ChatGPT setting. The initial rollout is through the OpenAI API, with access limited to a small group of customers as the company studies where the performance difference actually creates value. OpenAI says access will expand as capacity grows, but it has not announced consumer availability, a public launch date, or pricing.

That makes the title “OpenAI testing Ultrafast mode for ChatGPT” slightly more complicated than it first appears. It is best understood as technology that could eventually shape ChatGPT experiences, but is presently being trialed in business and developer products built on OpenAI’s platform.

The company’s early focus is practical rather than flashy: coding tools, customer support, commerce, financial research, incident response, and other interactive applications. These are settings where an answer that arrives a few seconds earlier can alter the workflow itself, not merely make it feel nicer.

Take incident response. When a production system breaks, engineers may need to sift through logs, traces, recent code changes, internal messages, and monitoring alerts while the issue is still unfolding. OpenAI says its own teams are using the mode to accelerate that loop: gathering evidence, proposing next checks, and helping prepare or validate a fix. The final decision and deployment still sit with humans, but a model that can keep pace with the investigation may reduce the idle gaps between each step.

Customer support offers another clear example. A voice or chat agent that pauses for several seconds while looking up policy details, stock availability, account history, or troubleshooting instructions makes the experience feel automated in the least flattering way. A high-speed model could make a more sophisticated assistant viable during a live conversation, rather than forcing companies to choose between a quick but simplistic bot and a better but slower one.

Why 750 tokens matter

“Tokens per second” is a technical metric, but its real-world meaning is reasonably straightforward. Tokens are the small chunks of text a model uses to generate language. Faster token output means the answer appears much more quickly and long responses can be completed with less waiting.

At OpenAI’s stated peak of 750 output tokens per second, a model could theoretically stream a substantial answer almost instantly. In practice, the experience will still depend on factors such as prompt size, tool calls, network latency, model reasoning, output length, traffic levels, and the exact application design. “Up to” is doing real work in the announcement; it is a peak claim, not a promise that every request will hit the same speed.

Still, the central idea is compelling. The biggest gains may not show up in a one-off prompt asking for an email draft. They are more likely to appear in repeated loops: a developer fixing a bug, an analyst testing a hypothesis, a researcher refining a search, or a support representative solving a messy problem across several systems.

OpenAI describes one research workflow in which teams might previously launch a batch of experiments overnight and check results the next morning. With much lower latency, the company argues that the same work can become an interactive cycle during the day: run an experiment, inspect the result, revise the approach, and run another.

That is a more consequential promise than “AI answers faster.” It suggests a future where the boundary between thinking, testing, and acting becomes thinner. The useful model is not simply the one that gives the best answer eventually, but the one that can join the user’s working rhythm without constantly interrupting it.

Cerebras is central

The hardware partner behind the preview is Cerebras, a company known for specialized AI systems designed around wafer-scale computing. OpenAI says its Ultrafast mode is powered by Cerebras as part of a partnership focused on ultra-low-latency inference – the process of running a trained model to generate an answer.

This is a meaningful detail because it shows that the next stage of AI competition is not only about model training. Training the next giant model remains expensive and strategically important, but serving that model efficiently at scale is equally critical. A highly capable model that is too slow, too costly, or too constrained to deploy broadly can struggle to become a genuinely useful product.

Cerebras says the implementation preserves GPT-5.6 Sol’s intelligence while delivering the speed improvement, although such claims should be judged over time through independent testing and real-world customer use. For now, OpenAI is explicitly treating Ultrafast as a preview, which is an acknowledgement that performance, capacity, reliability, and product design all need to be proven beyond a launch-day benchmark.

A shift in what AI can do

It is tempting to see Ultrafast as a simple infrastructure upgrade, the kind of technical change most people never notice. But speed can alter what is economically and ergonomically possible.

For a shopping assistant, latency may determine whether a customer stays engaged long enough to get an answer about fit, availability, delivery, or returns. For fraud detection, the value of a signal may fall quickly as a transaction progresses. For a developer tool, a slow assistant encourages context switching; a fast one may become part of the actual editing and debugging loop. OpenAI identifies commerce, financial research, security, coding, support, and live experimentation as the areas it is watching most closely during the trial.

The company is also testing the technology with an initial group of customers including Jane Street, Podium, Basis, and Rogo. John Crepezzi of Jane Street said the speed made it practical for developers to work “in a more focused and productive way” alongside models, which captures the broader thesis behind the launch: reducing delay can change behavior, not just benchmark charts.

For ChatGPT users, the immediate takeaway is more modest. There is no public Ultrafast button to enable, and no guarantee the exact tier will arrive in the consumer product. But OpenAI’s decision to test it offers a strong hint about where the company sees the next product advantage: not only smarter answers, but AI that can respond quickly enough to feel present in the moment.

If that idea holds up outside the limited preview, the real breakthrough may not be a chatbot that types faster. It may be a shift toward AI systems that are capable enough for difficult work and responsive enough that people actually want to keep them open while they do it.


Discover more from GadgetBond

Subscribe to get the latest posts sent to your email.

Topic:ChatGPT
Leave a Comment

Leave a ReplyCancel reply

Most Popular

Meet the Samsung Galaxy Card: perks, metal build, and rewards

Porsche launches Advanced Package for the Macan, Macan 4 and 4S

Notion now lets you pick your AI model

How Amazon Locker works for apartments, dorms, and busy schedules

Porsche crafts one-off Flachbau RS in tribute to ‘Moby Dick’

Also Read
A Dell laptop with the Windows logo displayed on its screen is shown on a colorful background with pink on top and blue on the bottom, viewed at an angle with part of the keyboard visible.

Microsoft overhauls the Windows 11 right-click menu

Samsung Galaxy Z Fold8 showing the Warranty and care hub with limited warranty details, Samsung Care+ eligibility, and device support categories.

Samsung’s One UI 9 brings diagnostics, repairs, and Care+ into one place

ElevenLabs MCP and Claude logos displayed side by side on a dark background, separated by a horizontal line.

ElevenLabs MCP arrives in Claude, bringing voice-agent control into the chat window

Futuristic digital artwork showing a glowing computer face icon inside a translucent glass-like sphere resting on a soft grassy surface. Floating reflective droplets surround the sphere against a dark black background, creating a surreal and minimalist sci-fi atmosphere.

Perplexity Computer adds Allow, Always Ask, Deny for connectors

A lineup of three black LG xboom Power party speakers in different sizes, each with circular illuminated woofers and the xboom logo on top, shown against a white background.

LG launches xboom Power 9000, 7000 and 5000 with AI karaoke tools

Porsche 911 GT3 100 Jahre Nürburgring

Green Hell tribute: Porsche builds a Nürburgring-only 911 GT3

White BMW ALPINA logo on a dark navy background, featuring the brand name around a circular emblem with stylized throttle bodies and a crankshaft.

BMW ALPINA relaunch

Red McLaren McL 6GT concept car parked outside a modern stone-and-glass garage.

McLaren’s McL 6GT brings back the manual supercar

Company Info
  • Homepage
  • Support my work
  • Latest stories
  • Company updates
  • GDB Recommends
  • Daily newsletters
  • About us
  • Contact us
  • Write for us
  • Editorial guidelines
Legal
  • Privacy Policy
  • Cookies Policy
  • Terms & Conditions
  • DMCA
  • Disclaimer
  • Accessibility Policy
  • Security Policy
  • Do Not Sell or Share My Personal Information
Socials
Follow US

Disclosure: We love the products we feature and hope you’ll love them too. If you purchase through a link on our site, we may receive compensation at no additional cost to you. Read our ethics statement. Please note that pricing and availability are subject to change.

Copyright © 2026 GadgetBond. All Rights Reserved. Use of this site constitutes acceptance of our Terms of Use and Privacy Policy | Do Not Sell/Share My Personal Information.