Anthropic is pushing its flagship Claude line forward again, and this time the headline isn’t simply “more capable.” Claude Opus 5.5 is designed to make demanding AI work cheaper, faster, and more practical for long-running agents.
Anthropic introduced Claude Opus 5.5 on September 22 as the first model in its new Claude 5.5 family. The company says the model performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Claude Opus 5 on typical workloads.
That combination matters because Anthropic is increasingly positioning Claude not just as a chatbot, but as an agent capable of spending hours working through software projects, research tasks, business workflows, and other complicated jobs.
Claude Opus 5.5 is built for long-running work
Anthropic describes Opus 5.5 as a major step up from Opus 5, particularly for tasks that require an AI agent to keep working through a large problem rather than answering a single prompt.
The company says an early tester used Opus 5.5 to complete a migration involving 680,000 lines of code in less than a day. In another test, the model successfully found ways to reduce loading times across a web application 39 out of 40 times, while avoiding some of the behavioral changes introduced by Opus 5 during its attempts.
Those examples point toward the direction Anthropic is taking with the Opus line: less emphasis on producing an impressive answer in a single interaction and more emphasis on getting an actual job finished.
Coding is particularly important here. Anthropic says Opus 5.5 is designed for sprawling software tasks such as codebase migrations, audits, debugging, and multi-file changes.
In one internal comparison, the company had Opus 5.5 and Fable 5.1 translate HAProxy from C to Rust. Both models produced implementations that passed nearly all of HAProxy’s regression tests, but Opus 5.5 completed the task in 9.5 hours compared with 12 hours for Fable 5.1, while costing 51% less.
Anthropic also says an early tester audited and fixed a 200,000-line codebase with Opus 5.5 in under three hours, compared with more than 20 hours for Opus 5.
The company’s benchmark results tell a similar story. On Terminal-Bench 4.0, Opus 5.5 scored 66.4%, compared with 52.3% for Opus 5. On FrontierCode v1.1, it scored 54.4%, while Opus 5 scored 48%. CursorBench 4.0 produced a 57.8% score for Opus 5.5 versus 46.6% for Opus 5.
Anthropic cautions that benchmark differences at this level do not always translate directly into equivalent differences in real-world use. The company also notes that its results were collected with its production safeguards enabled, which can intervene in some high-risk tasks.
The bigger story is efficiency
The performance improvements would be notable on their own, but Anthropic is putting just as much emphasis on how much it costs to achieve them.
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens. Opus 5 cost $5 per million input tokens and $25 per million output tokens.
Cache reads are also considerably cheaper, falling to $0.20 per million tokens from $0.50. Anthropic says cache reads account for a large share of the cost of long-running agentic and coding workloads.
The company estimates that Opus 5.5 costs about 40% less to run than Opus 5 on typical workloads when differences in token usage and compute are taken into account. It also says the new model generates output more than 30% faster.
For developers who want even more speed, Anthropic is offering a Fast mode for Opus 5.5 in Claude Code and on the Claude Platform. Fast mode can provide up to 2.5 times the speed, but at $8 per million input tokens and $40 per million output tokens.
That makes Opus 5.5 particularly interesting for agentic workloads, where the cost isn’t determined only by how expensive a single model response is. An agent may make dozens or hundreds of model calls while reading files, running commands, checking its work, fixing mistakes, and continuing toward a goal.
Reducing both token consumption and the number of steps can therefore have a much bigger impact than a simple per-token price reduction suggests.
Anthropic is also changing how Claude communicates
Opus 5.5 isn’t only about coding benchmarks.
Anthropic says it has made significant changes to the model’s communication style based partly on feedback from users of Opus 5. The company says Opus 5.5 puts important information up front, uses less jargon and fewer idiosyncratic phrases, and follows user-provided writing instructions more consistently.
That may sound like a relatively small improvement compared with a benchmark jump, but it becomes more important when the model is acting as a collaborator over several hours.
Anthropic’s own examples show Opus 5.5 producing shorter, more direct explanations of software bugs while retaining the important technical details. The company argues that clearer communication also makes agent output easier for people to inspect and verify.
Several companies that tested the model reported similar improvements.
GitHub says Opus 5.5 used among the fewest tokens and steps it measured across its testing in GitHub Copilot CLI and VS Code. Clio reported running a large engineering task across six repositories for more than 18 hours with the model working unattended. Lovable said its testing showed the model completing tasks in roughly one-third to one-half fewer steps in some workflows.
Those are company-reported evaluations rather than independent benchmarks, so they should be viewed in that context. Still, they illustrate the type of workflow Anthropic is targeting.
Claude Opus 5.5 gets stronger safety safeguards
As models become more capable of operating autonomously, the question isn’t only whether they can complete a task. It’s whether they can do so without taking actions they shouldn’t.
Anthropic says Opus 5.5 performed better than its recent models on its automated behavioral audit, which covers nearly 2,000 simulated scenarios. The company says the model was less likely to take difficult-to-reverse actions or operate outside its assigned boundaries, while also showing stronger resistance to prompt injection than Opus 5.
The accompanying system card documents the model’s safety evaluations and deployment decisions.
Opus 5.5 also launches with safeguards similar to those used for Anthropic’s Fable 5.1 model in areas including cybersecurity, biology, and model distillation. For certain cybersecurity tasks, Anthropic says requests can be routed to Claude Opus 4.8 instead.
Anthropic is also using an action-screening classifier for its coding agent, along with an open-source sandbox that security teams can audit and code-review mechanisms designed to catch vulnerabilities before changes are merged.
The company says Opus 5.5 matched or exceeded Opus 5 against the prompt-injection attacks it tested across coding, tool use, computer use, and web browsing.
For API customers, Opus 5.5 also retains the preserved-thinking mechanism introduced with Fable 5.1. Anthropic says the feature is intended to make it harder for attackers to extract a model’s capabilities through large-scale distillation attacks.
Opus 5.5 is already available
Claude Opus 5.5 is available to Pro, Max, Team, and Enterprise users through Claude. Developers can access it through the Claude Platform, as well as Amazon Web Services, Google Cloud, and Microsoft Foundry.
Anthropic is also increasing five-hour usage limits for Pro, Max, Team, and seat-based Enterprise plans. Subscription users get a rate-limit reset that can be saved and used when needed.
For developers, the API model identifier is claude-opus-5-5.
Anthropic says Opus 5.5 can also be used with zero data retention, while the model includes watermarking measures designed to comply with the EU AI Act. Thinking mode can no longer be disabled for the model.
The company isn’t stopping at Opus, either. Claude Sonnet 5.5 and Claude Haiku 5.5 are expected to arrive in the coming weeks, bringing many of the same improvements in performance, efficiency, and safety to the other tiers of the Claude lineup.
For now, though, Opus 5.5 represents an interesting shift in the AI model race. The pitch isn’t simply that Claude can solve harder problems. Anthropic is arguing that it can solve those problems with fewer steps, fewer tokens, less time, and less money.
For anyone actually putting AI agents to work on large software projects or professional workflows, that may ultimately matter more than another few points on a benchmark.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
