The lineup now splits by job: the Sonnet tier for scoped, everyday work and Opus for planning and review. The benchmark figures circulating come from a secondhand summary, so treat them as claims until Anthropic's own write-up is checked, and measure the two models on your own tasks before changing anything.
Claude Sonnet 5.5 is out! We wrote a guide for building with it: • choosing between Sonnet 5.5 and Opus 5.5 • migrating from Sonnet 5 and tuning effort • using it in Claude Code https://t.co/YAQBD95pF8 https://t.co/exQKLA9rev
— ClaudeDevs (@ClaudeDevs) · View on X
- What happened: On September 29, @ClaudeDevs (https://x.com/ClaudeDevs/status/2104687805876367793) announced Claude Sonnet 5.5 with a guide covering when to pick Sonnet 5.5 over Opus 5.5, how to migrate from Sonnet 5, and how to tune effort. The same day, @github (https://x.com/github/status/2104637226336538862) made it generally available in Copilot, describing it as built for "well-scoped everyday work like building features and fixing bugs" and reporting that it matched Sonnet 5 on coding tasks with fewer steps, tokens and tool calls.
- The numbers in circulation: @Israfilv2 (https://x.com/Israfilv2/status/2104637819402764747) summarised the release as over 30% faster than Sonnet 5, up to 30% fewer tokens on most work, 80.9% on SWE-bench Verified, 93.5% on Terminal-Bench 2.0, and unchanged pricing at $3 per million input tokens and $15 per million output. No post from Anthropic with those exact figures appears in the sources read for this piece, so verify before quoting.
- Why it matters this week: It follows Opus 5.5, which @DamiDefi (https://x.com/DamiDefi/status/2103117499621625951) reported at $4 per million input and $20 per million output, with typical workloads costing 40% less than Opus 5. @rauchg (https://x.com/rauchg/status/2103216656747262419) shared Vercel AI Gateway data showing Anthropic's share of spend falling from 69% to 40% over two months while OpenAI rose from 10% to 24%, with Opus 5.5 reaching 10% of spend within two days of release. Two efficiency-focused releases in one week read as a response to that pressure.
- Routers are changing who picks the model: @kimmonismus (https://x.com/kimmonismus/status/2104628247296409952) described a Jev-based router that reportedly reaches 99% of Opus 5.5 and GPT-6 Astra performance on three benchmarks for 40% less. @gradientintern (https://x.com/gradientintern/status/2103363088397250605) reported a simple router hitting 75 of 100 on SWE-bench Verified at $25.66 against $54.73 for always using Opus. Both are self-reported results.
- What to do this week: Set Sonnet 5.5 as the default for scoped tasks and keep Opus 5.5 for planning and code review, the split @swapnilskr (https://x.com/swapnilskr/status/2103844926479822852) already runs with Opus as pilot and cheaper models on pre-designed work. Rerun one real task on both models and record tokens, time and whether tests passed. @databricks (https://x.com/databricks/status/2103207342061883611) notes a new model arrives roughly every five days, so put the model choice in one shared configuration instead of each developer's habits.
Which Claude model should be the default now?
Sonnet 5.5 for anything with a clear scope and a test to pass, Opus 5.5 for planning, architecture calls and reviewing other agents' output. That is how GitHub positioned the release and how @Israfilv2 (https://x.com/Israfilv2/status/2104637819402764747) summarised Anthropic's own guidance: use Sonnet first for daily work and keep Opus for jobs that need deeper judgment.
Should I switch today or wait for a router to decide?
Switch the default today, because the price did not change and the efficiency claim is easy to test on your own repo. Routers such as the ones @kimmonismus and @gradientintern describe show the direction: the model choice becomes a per-step decision made by software. Until you run one, a two-tier default with a weekly review of spend gets most of the benefit.
What breaks if I change models without retesting?
Prompts tuned for the previous model. Anthropic's own Opus 5.5 guidance, as reported by Search Engine Journal (https://www.searchenginejournal.com/anthropic-claude-opus-5-5-prompting-guidance/591278/), asks developers to retest effort settings and reconsider prompt lines that tell the model to think carefully. @mika_systems (https://x.com/mika_systems/status/2103921273852108955) found Opus 5.5 at maximum effort spent 64,899 thinking tokens on six coding tasks where a decision layer passed the same checks with 954, so effort settings alone can move cost more than the model swap.
What does this change for a Shopify team building pages?
Little on the storefront, more on the bill. If an agency or in-house team uses Claude Code or Copilot to write theme sections and product copy, the everyday tier just got faster at the same price, and the expensive tier should be reserved for reviewing what ships. @ClaudeDevs (https://x.com/ClaudeDevs/status/2104676099083190435) also published guidance on building evaluations that Claude Code can improve against, which is the missing piece for comparing two models on a real store task.
Sources
- @ClaudeDevs: https://x.com/ClaudeDevs/status/2104687805876367793
- @github: https://x.com/github/status/2104637226336538862
- @Israfilv2: https://x.com/Israfilv2/status/2104637819402764747
- @DamiDefi: https://x.com/DamiDefi/status/2103117499621625951
- @rauchg: https://x.com/rauchg/status/2103216656747262419
- @kimmonismus: https://x.com/kimmonismus/status/2104628247296409952
- @gradientintern: https://x.com/gradientintern/status/2103363088397250605
- @swapnilskr: https://x.com/swapnilskr/status/2103844926479822852
- @databricks: https://x.com/databricks/status/2103207342061883611
- @mika_systems: https://x.com/mika_systems/status/2103921273852108955
- @ClaudeDevs: https://x.com/ClaudeDevs/status/2104676099083190435
- Search Engine Journal: https://www.searchenginejournal.com/anthropic-claude-opus-5-5-prompting-guidance/591278/