OpenAI is making its high-end AI models considerably cheaper to run. The company has launched GPT-6.1 Sol, an upgraded version of GPT-6 Sol that OpenAI says comes surprisingly close to its flagship GPT-6 Astra on several demanding workloads—while costing roughly one-fifth as much.
That makes GPT-6.1 Sol less about simply pushing model performance higher and more about making advanced AI practical to deploy at scale. For developers building coding agents, computer-use systems and other long-running workflows, the cost of every request can matter just as much as the model’s raw intelligence.
OpenAI describes GPT-6.1 Sol as a new balance between capability and cost. The model improves on GPT-6 Sol across coding, professional work, computer use, scientific research and factuality, with several of its evaluation results approaching GPT-6 Astra.
GPT-6.1 Sol gets much closer to Astra
The biggest improvement shows up in agentic coding. On DeepSWE v1.1, a benchmark focused on complex software-engineering tasks in real codebases, OpenAI says GPT-6.1 Sol matches GPT-6 Astra while costing roughly one-fifth as much. It also improves GPT-6 Sol’s best score by 6.4 percentage points while using a lower reasoning effort.
The model is also aimed at the kind of professional work where AI has to understand more than plain text. On GDP.pdf, which tests models on complex documents containing tables, charts, diagrams and fine-print details, OpenAI says GPT-6.1 Sol approaches Astra’s performance at around one-fifth of the cost per task.
For multi-step business workflows, GPT-6.1 Sol also posts a 4.8-percentage-point improvement over GPT-6 Sol at the same medium reasoning setting on AutomationBench. OpenAI says it scored 2.2 points above Opus 5.5 in that test while costing roughly one-third as much.
Computer use is another area where OpenAI is targeting more capable agents. On OSWorld 2.0’s offline set, GPT-6.1 Sol reportedly scores seven percentage points higher than GPT-6 Sol at maximum reasoning effort while costing less than half as much. Its score comes within 2.1 percentage points of Astra’s, according to OpenAI, at roughly one-seventh of the cost per task.
Scientific workloads show a similar pattern. On Terminal-Bench Science 0.1, OpenAI says GPT-6.1 Sol more than doubles GPT-6 Sol’s score at maximum reasoning effort while costing less than half as much per task. The company puts the average cost at $5.47 per task, compared with $23.21 for Opus 5.5 and $23.80 for Astra. Astra still posted the highest score in that particular evaluation, however.
OpenAI is also targeting factuality
GPT-6.1 Sol isn’t only about making agents cheaper. OpenAI says it also improves factual accuracy over GPT-6 Sol on a particularly difficult set of prompts drawn from conversations where users had previously flagged factual errors.
At low reasoning effort, the percentage of responses containing a factual error fell from 11.4% with GPT-6 Sol to 7.7% with GPT-6.1 Sol, according to OpenAI. That’s approximately a 32% reduction in the measured error rate.
There’s an important caveat here: OpenAI says these prompts were deliberately selected to trigger difficult factual failures and aren’t representative of typical ChatGPT usage. In other words, the numbers are useful for comparing the models in that evaluation, but shouldn’t be interpreted as a general real-world error rate.
OpenAI also says GPT-6.1 Sol improves on its alignment evaluations, including being more transparent about limitations, following explicit restrictions and avoiding unauthorized outcomes during agentic tasks.
The pricing is the real story
The headline feature of GPT-6.1 Sol may ultimately be its price.
OpenAI is charging $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens through the API. The cached-input price is particularly notable because OpenAI says it is 95% lower than standard input pricing and 50% lower than GPT-6 Sol’s cached-input price.
That’s important for AI agents that repeatedly work with the same context. Instead of paying the full input cost every time an agent returns to a large codebase, document set or workflow, developers can make greater use of cached context.
The broader strategy is easy to see: make a model that is close enough to the frontier for many workloads, then make it inexpensive enough that developers can actually run it frequently.
That’s a very different proposition from simply having the most capable model available.
GPT-6.1 Sol isn’t replacing Astra
Despite the performance gains, GPT-6.1 Sol isn’t positioned as a complete replacement for GPT-6 Astra.
OpenAI still says Astra has the highest score on its Terminal-Bench Science evaluation and recommends it for the most difficult scientific research tasks. GPT-6.1 Sol is instead positioned as the model for workloads where developers want much of Astra’s capability without paying Astra-level prices.
That distinction could become increasingly important as AI agents move from occasional assistants to systems that perform hundreds or thousands of model calls. A small difference in capability can matter less than a huge difference in operating cost when an agent has to repeatedly reason, use tools and inspect results.
GPT-6.1 Sol is available now, but not in regular ChatGPT
GPT-6.1 Sol is available starting September 29 to Plus, Pro, Business, Enterprise and Edu users through ChatGPT Work and Codex. It is not yet available in regular ChatGPT.
Developers can access the model through the OpenAI API using the gpt-6.1-sol model name. OpenAI also says a GPT-6.1 Sol Ultrafast option is coming to Codex in the next few days, with up to eight times faster token generation than the standard speed.
For OpenAI, then, GPT-6.1 Sol is less of a traditional generational leap and more of a scaling play. The company already has a flagship model at the top of its lineup. Now it wants to bring a large portion of that capability down to a price where developers can use it much more aggressively.
And if OpenAI’s benchmark results translate into real-world agent performance, that could be just as important as another jump at the very top of the model hierarchy.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
