GitHub's Copilot note says the new model matches the old one with fewer steps and tokens, but the week's published cost tests point to routing, effort settings and context as the larger levers on a team's monthly bill.

Switch the everyday model first because it costs nothing to try, then work down the ladder: a router that picks a cheaper model per step, a decision layer that caps thinking, and a cleaner harness. Every figure below is self-reported by the team that ran it, and none is a monthly invoice from an agency.

  • What happened: On September 29, @github (https://x.com/github/status/2104637226336538862) made Claude Sonnet 5.5 generally available in Copilot and said that in its early testing the model matched Sonnet 5 on coding tasks while using significantly fewer steps, tokens and tool calls, and finished faster. @Israfilv2 (https://x.com/Israfilv2/status/2104637819402764747) summarised the release as up to 30% fewer tokens on most work at an unchanged $3 per million input and $15 per million output tokens. For teams billed per token that is a lower cost per task; for teams on seat plans with usage limits, the bill stays flat and the limit stretches.
  • The cost ladder this week's tests describe: @gradientintern (https://x.com/gradientintern/status/2103363088397250605) reported a simple router reaching 75 of 100 on SWE-bench Verified at $25.66 against $54.73 for always using Opus, a 53% cut. @mika_systems (https://x.com/mika_systems/status/2103921273852108955) ran six coding tasks twice: Opus 5.5 at maximum effort spent 64,899 thinking tokens, a decision layer in front of it spent 954, both passed 12 of 12 checks, and the API total fell from $3.76 to $1.70. @kimmonismus (https://x.com/kimmonismus/status/2104628247296409952) relayed a reported 99% of Opus 5.5 and GPT-6 Astra performance for 40% less via a Jev-based router. @jamiegrove (https://x.com/jamiegrove/status/2102858267441430708) proofread 682 Shopify product pages with the same model and prompt under two harnesses: the first burned about $185 and never finished, the second cost $12.60.
  • Why the invoice can still climb: @jensenloke (https://x.com/jensenloke/status/2104252961430102289) described a five-day Devin session on ops cleanup that consumed about 688 million tokens of context, reading roughly 129 tokens for every one written. @cyrilxbt (https://x.com/cyrilXBT/status/2103752714941538621) counted five or six separate APIs behind one agent, for embedding, reranking, PDF parsing, safety and reasoning, each with its own bill. @rauchg (https://x.com/rauchg/status/2103216656747262419) shared Vercel AI Gateway data in which Anthropic's share of spend fell from 69% to 40% over two months while OpenAI's rose from 10% to 24%, a shift between vendors rather than a drop in total.
  • What to do this week: Make Sonnet 5.5 the default for scoped tasks and retest effort settings, since @mika_systems' numbers show effort can move cost more than the model. Cap session length so context stops compounding. Put model choice, routing and spend limits in one shared configuration, the approach @databricks (https://x.com/databricks/status/2103207342061883611) describes for teams facing a new model roughly every five days. Record spend per session before and after, using a local tool such as the one @DataChaz (https://x.com/DataChaz/status/2096303020653101200) highlighted, so the comparison is your own number.

Which lever should a team pull first?

The free one: set the cheaper model as default for well-scoped work, which is how @github positioned Sonnet 5.5. Then fix effort settings, because the largest single saving in this week's tests came from cutting thinking tokens, not from changing models. Routers and harness rebuilds come after, since they take engineering time and their published gains are self-reported.

Why can the monthly bill rise while each task gets cheaper?

Because tasks are a small share of tokens. Context reads dominated the session @jensenloke measured, agents call several paid services per step according to @cyrilxbt, and cheaper per-task pricing tends to invite more runs. @rauchg's gateway data shows spend moving between providers within two months without shrinking, which is what a team sees when it swaps models without a budget cap.

How does a team know the lever worked?

By holding the task fixed and measuring three things GitHub measured: steps, tokens and tool calls, plus time and whether tests passed. @mika_systems' setup is the template: same six tasks, two configurations, hidden checks, and a dollar total per run. Anthropic's own guidance, as reported by Search Engine Journal (https://www.searchenginejournal.com/anthropic-claude-opus-5-5-prompting-guidance/591278/), asks developers to retest effort when a model changes, so keep effort constant between the two runs.

Sources