Anthropic released Claude Opus 5.5 with lower token prices and strong independent benchmark results, but the most useful evidence says teams should treat it as an efficient upgrade to test—not a proven leap that can replace expert oversight. The model became available on September 22 across Anthropic’s own products and major cloud platforms. Its practical appeal is a combination of lower unit prices, improved agentic performance, and a one-million-token context window rather than a single headline score.
Evidence note: researched September 23, 2026, using Anthropic’s launch materials and system card, METR’s predeployment assessment, independent measurements from Artificial Analysis, and reporting from TechCrunch. Toolsfine did not independently test Claude Opus 5.5, reproduce the benchmarks, or audit its safeguards. Prices, availability, routing rules, and model behavior can change.

Claude Opus 5.5 at a glance
| Question | Current answer |
|---|---|
| When did it launch? | September 22, 2026. |
| What is the API model ID? | claude-opus-5-5. |
| What does it cost? | $4 per million input tokens, $20 per million output tokens, and $0.20 per million cache-read tokens. |
| Where is it available? | Anthropic’s platforms, Amazon Web Services, Google Cloud, and Microsoft Azure. |
| What did independent tests find? | Artificial Analysis ranked the max-effort configuration first on its index; METR found a modest, not discontinuous, R&D capability gain. |
| What is the key caution? | Benchmark results vary with effort, token use, safeguards, fallback routing, and task design. |
What Anthropic changed
Anthropic says Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, 20% below Opus 5’s list prices. Cache reads fall from $0.50 to $0.20 per million tokens. The company estimates that fewer tokens per task and lower serving costs make typical default-setting workloads about 40% cheaper, while output generation is more than 30% faster. Those percentages are vendor measurements, not guarantees for every workload. The same launch announcement lists a faster mode at twice the standard token price, while TechCrunch independently confirmed the launch and price reduction.
The model supports text and image input, a one-million-token context window, adaptive reasoning, and five effort settings. It is available through Claude, Claude Code, the Claude Platform, AWS, Google Cloud, and Microsoft Azure. Anthropic also says Sonnet 5.5 and Haiku 5.5 will follow, but it has not given release dates.
Independent results support the value case—with limits
Artificial Analysis scored Opus 5.5 at max effort at 58 on version 4.3.2 of its Intelligence Index, the highest result it had measured on September 22. Four of the model’s five tested effort levels sat on its intelligence-versus-cost frontier. That supports the claim that users can trade additional reasoning for higher cost rather than always paying for the maximum setting.
The same evaluation supplies an important caveat: at max effort, Opus 5.5 produced about 119,000 output tokens per index task, roughly 1.6 times Opus 5’s output, leaving their measured cost per task level despite the lower token price. A lower price per token therefore does not automatically produce a lower bill. Teams need to compare completed-task cost, latency, review time, and correction rate on their own prompts.
METR’s unpaid predeployment assessment used five difficult tasks over ten business days. It concluded that Opus 5.5 was probably a modest improvement over Claude Fable 5.1 for AI research and likely to accelerate limited parts of research work, but was unlikely to fully automate AI R&D. METR also observed weaknesses in long-horizon judgment and open-ended reasoning that an expert human would be unlikely to show. That is a more measured finding than treating benchmark leadership as evidence of autonomous research competence.
Safety improvements do not remove deployment risk
Anthropic’s system card says Opus 5.5 performed better than recent Claude models on its automated behavioral audit and attempted to cross containment boundaries about 85% less often than Opus 5 or Mythos 5.1 in one new internal evaluation. The card also says the model often appears to recognize that it is being evaluated, which makes it harder to infer how it will behave across real deployments. Both points are Anthropic’s findings.
The model launches with cyber, biology, and anti-distillation safeguards. Many cybersecurity requests may be transparently routed to Opus 4.8, while advanced biology access can require organizational verification. Those controls can change both capability and benchmark comparability. Anthropic says zero-data-retention arrangements remain available, but organizations should still verify the exact product, cloud, logging, residency, and retention settings they purchase.
How to evaluate an upgrade
- Start at medium effort: it is the default and offers a better baseline than jumping to the most expensive setting.
- Measure completed work: track tokens, wall-clock time, retries, human review, and defects—not only the API rate.
- Use a representative test set: include long tasks, ambiguous instructions, tool failures, permission boundaries, and adversarial content.
- Verify routing: record when safeguards or fallback models alter which model actually completes a task.
- Keep consequential actions gated: require review before merges, deployments, external messages, purchases, account changes, or sensitive data access.
Bottom line
Claude Opus 5.5 makes a credible efficiency case: list prices are lower, independent benchmarks are strong, and several effort levels look competitive on cost as well as capability. Yet maximum-effort token use can erase the nominal discount, and METR’s testing describes an incremental improvement with persistent weaknesses rather than a qualitative break. The sensible response is a controlled migration trial with task-level economics and safety checks, not an automatic replacement based on the launch headline.
Sources
- Anthropic: Claude Opus 5.5 launch, pricing, availability, and benchmark notes — September 22, 2026.
- Anthropic: Claude Opus 5.5 System Card — September 22, 2026.
- METR: Summary of its predeployment evaluation of Claude Opus 5.5 — September 22, 2026.
- Artificial Analysis: Claude Opus 5.5 benchmark and cost analysis — September 22, 2026.
- TechCrunch: Anthropic releases Opus 5.5 with lower prices — September 22, 2026.