Anthropic released Claude Opus 5.5 this week, cutting prices on its flagship model by roughly 40% on typical workloads while posting the strongest agentic coding scores the company has published to date.
The new pricing sets input tokens at $4 per million and output tokens at $20 per million, each about 20% below Opus 5. Cached reads fall even further, down 60% to $0.20 per million tokens, and a new "fast mode" is available at $8 input and $40 output per million tokens for workloads that prioritize speed over depth. Anthropic said the combined effect of the pricing changes and efficiency gains amounts to a 40% cost reduction for customers running comparable jobs.
Wider benchmark gaps
Anthropic's released benchmark figures show Opus 5.5 pulling ahead of both its own predecessor and competing frontier models on tasks that involve extended, multi-step work rather than single-turn question answering.
On Terminal-Bench 4.0, a test of agentic coding in real command-line environments, Opus 5.5 scored 66.4%, compared with 55.8% for a rival model Anthropic benchmarked against and 52.3% for Opus 5 itself. On FrontierCode v1.1 it reached 54.4%, and on CursorBench 4.0 it scored 57.8%, both meaningful jumps over the prior generation. On GDPval-AA, a benchmark built around real-world knowledge work rather than synthetic tests, Opus 5.5 posted an Elo rating of 1846, again ahead of both comparison points.
Anthropic also highlighted a 30% improvement in output generation speed and pointed to customer accounts of the model completing longer unsupervised runs. A staff engineer at analytics company Clio said the model "stayed on task for over 18 hours" on one project and "required minimal reworking" once it finished. An executive manager at Quantium described a complex coding task that previously took 38 prompts spread over four days coming in at 11 prompts over three hours with the new model. GitHub's chief product officer said Opus 5.5 "used among the fewest tokens and steps measured" across its internal testing.
What changed under the hood
Beyond raw benchmark scores, Anthropic pointed to two specific behavioral changes. The model is now 85% less likely to attempt to circumvent guardrails it encounters during a task, according to the company's internal testing, and its resistance to prompt injection, attempts by external content to hijack a model's instructions, matches or exceeds Opus 5 despite the model's added capability.
Anthropic said Opus 5.5 also ships with additional safeguards specific to cybersecurity and biology-related work, a watermarking feature intended to support compliance with the EU AI Act, and a measure the company calls "preserved thinking" designed to make the model's internal reasoning harder to extract and reuse for training competing systems.
Independent safety review before release
Before shipping, Anthropic had the model evaluated by two outside organizations, METR and Frontier Design, running an automated behavioral audit across nearly 2,000 scenarios designed to probe for unsafe or deceptive behavior. Anthropic said Opus 5.5 posted the highest scores of any model it has released on that audit.
The company has increasingly leaned on third-party evaluation as agentic models are given more autonomy to act across multiple steps without a human reviewing each one, a shift that has drawn scrutiny from AI safety researchers and, this week, from the United Nations Security Council, which held its first formal briefing on AI and international security on September 23.
Why the pricing move matters
The price cut lands as enterprise customers increasingly run AI agents continuously rather than in short bursts, a shift that makes per-token cost a much larger line item than it was for chatbot-style usage. A head of quantitative research at investment firm Walleye Capital said of Opus 5.5's analytical output: "No model we've tested had caught and acted on that before," pointing to the model's use in financial research workflows where cost per query compounds quickly at scale.
Anthropic's move follows a broader industry pattern in 2026 of frontier labs cutting per-token prices even as model capability climbs, a trend driven partly by improved inference efficiency and partly by intensifying competition for enterprise agentic-AI workloads.
Sources
- Anthropic — Introducing Claude Opus 5.5, September 2026.

