, , , ,

Anthropic Cut the Price Where Agents Actually Spend: the Cache

Close-up detail of coiled server power cables and amber cache-drive LEDs behind a scuffed rack bezel, dust on the grille, cool blue aisle li

Anthropic launched Claude Opus 5.5 on September 22, the first model in its 5.5 family, promising performance it compared to Fable 5.1 on most tasks at roughly 40 percent lower cost than Opus 5 — a cut aimed at the coding agents that now burn tokens in the cache.

List prices fell to $4 per million input tokens and $20 per million output tokens, each 20 percent below Opus 5. Cache reads dropped 60 percent, to $0.20 per million, which is where agentic and coding workloads actually spend. A faster mode is offered at $8 and $40 per million. The company said the model generates output more than 30 percent faster than Opus 5 and is live on Claude, AWS Bedrock, Google Cloud, and Microsoft Azure under the API name claude-opus-5-5.

Sonnet 5.5 and Haiku 5.5 are due in the coming weeks. Anthropic said it ran stronger safety evaluations with external testers including METR and Frontier Design, with safeguards described as similar to those on Fable 5.1.

The last price war, including OpenAI's August cut on GPT-5.6 Sol, was about headline tokens. This one is about memory. If cache is 60 percent cheaper, the total cost of an overnight coding agent moves more than the brochure rate. That is a shot at whoever is billing verbose tool loops. It is also a bet that enterprises will switch models for a spreadsheet, not a benchmark chart.

Image source: i.ibb.co