claude sonnet 5.5 vs opus 5.5: which should a small business use for everyday work?

elisabeth hitz · published september 29, 2026 · updated september 29, 2026 · 6 min read

anthropic released claude sonnet 5.5 on september 28, 2026, six days after opus 5.5. on paper they now do almost the same knowledge work at very different prices. that is a routing decision, and most small businesses will make it by habit instead of on purpose.

for most everyday small business work, use claude sonnet 5.5 and keep opus 5.5 for the jobs where a subtle mistake is expensive. sonnet 5.5 scores nearly level with opus 5.5 on real-world knowledge work at half the per-token price. anthropic still positions opus for careful judgment, so route by task shape, not benchmark.

the near-tie is real. it is also one benchmark, and the difference between "nearly level on bounded tasks" and "equal" is exactly where a business loses money. so here is what actually changed, what the numbers do and do not say, and a four-step test to run on your own work before you move anything.

what anthropic actually released

sonnet 5.5 is the mid-tier claude model, released on september 28, 2026 as a faster, lower-cost complement to opus 5.5. it is available in the claude apps, on the claude platform as claude-sonnet-5-5, and through amazon web services, google cloud, and microsoft azure (anthropic, sep 28 2026). a smaller haiku 5.5 is described as coming in the following weeks, not released (anthropic, sep 28 2026).

anthropic draws the line between the two models in one sentence. opus 5.5 is built for complex work that needs careful judgment, while sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and polished documents, slides, and spreadsheets (anthropic, sep 28 2026). that sentence is more useful to a business owner than any benchmark in the post.

the numbers, side by side

on list price, sonnet 5.5 costs half of opus 5.5, and on one knowledge work benchmark they score two points apart. every figure below is from anthropic's announcements, with the gdpval-aa scores as reported by venturebeat.

what you are comparingclaude sonnet 5.5claude opus 5.5
releasedseptember 28, 2026 (anthropic)september 22, 2026 (anthropic)
input price, per million tokens$2 (anthropic, sep 28 2026)$4 (anthropic, sep 22 2026)
output price, per million tokens$10 (anthropic, sep 28 2026)$20 (anthropic, sep 22 2026)
gdpval-aa, real-world tasks across 44 occupations1,844 (venturebeat, sep 28 2026)1,846 (venturebeat, sep 28 2026)
anthropic's positioningwell-scoped everyday tasks, bug fixes, documents, slides, spreadsheetscomplex work requiring careful judgment, long autonomous runs
send itdrafts, summaries, formatted documents, clearly defined tasksambiguous, high-stakes, or long multi-step jobs

three savings claims that are easy to mix up

there are three different cost claims floating around this week, and each one compares against a different thing. keep them apart or you will budget off the wrong one.

  1. half the price is sonnet 5.5 vs opus 5.5, per token. $2 and $10 against $4 and $20 per million tokens (anthropic, sep 22 and sep 28 2026).
  2. up to 30% less is sonnet 5.5 vs sonnet 5, per task. the per-token price did not change. anthropic says sonnet 5.5 usually gets the same work done with far fewer tokens (anthropic, sep 28 2026), and venturebeat's launch coverage attributes the saving to fewer tokens and fewer tool calls (venturebeat, sep 28 2026).
  3. 40% less is opus 5.5 vs opus 5, per workload. anthropic says opus 5.5 costs about 40% less than opus 5 to run on typical workloads (anthropic, sep 22 2026).

the second one is the trap. a saving that comes from fewer tokens per task never shows up on a price sheet. it only shows up if you measure what a finished job costs, and the price sheet is where most people stop. i ran the same who-pays-for-which comparison for claude opus 5 vs fable 5.

what the near-tie does not tell you

a two-point gap on gdpval-aa means the models are close on bounded knowledge work, not that they are interchangeable. gdpval-aa tests real-world tasks across 44 occupations and nine industries, and anthropic says sonnet 5.5 lands about 400 points above sonnet 5 on it (anthropic, sep 28 2026). that is a big jump for everyday work. it says little about the task that runs for an hour, changes direction halfway, and has to notice its own mistake.

that second kind of task is exactly what anthropic keeps opus for. if your business has a handful of those (a contract review, a pricing model, a messy client migration), they are where the extra spend buys you something. everything else is paying opus rates for sonnet work. the broader case for splitting work across models is in you are using one ai model for everything.

will it break what you already built?

early customer reports say switching improved results without prompt changes, but you should check on your own tasks before trusting that. slack's principal engineer put it plainly in anthropic's announcement:

"sonnet 5.5 did better than sonnet 5 on almost all of our offline slackbot evals"
curtis allen, principal engineer, slack, in anthropic's claude sonnet 5.5 announcement, sep 28 2026

the same quote adds that it got there in fewer steps and with about 14% fewer output tokens (anthropic, sep 28 2026). "almost all" is doing honest work in that sentence. slack has an eval suite. you probably have three real tasks from last week, and that is enough.

the four-step routing test

run this on your own work in about twenty minutes, and you will know which jobs move and which stay.

  1. write the task as one sentence. if you can describe the finished output in one sentence, it is well scoped, which is where anthropic says sonnet 5.5 is strongest.
  2. ask what a wrong answer costs. if a subtle error would cost money, a client, or a legal problem, and nobody would catch it, it is a careful-judgment task. it stays on opus.
  3. run three real examples on both. recent tasks from your own business, side by side. your work, not a leaderboard.
  4. compare cost per finished task. not per token. sonnet 5.5's advantage over sonnet 5 lives in fewer tokens and tool calls, so it only appears per task.

if you work in claude code, the effort setting is a second dial on the same decision, and i covered how to pick it in which claude effort level should you use. if you route across models automatically, read you might not be running the model you benchmarked first.

the takeaway

sonnet 5.5 moves the default. for drafting, documents, spreadsheets, and anything you can describe in one sentence, it now does close to opus-level work at half the per-token price, and it probably finishes in fewer tokens than the sonnet you were using. opus 5.5 is still the model for the few jobs where judgment is the product. the mistake is not picking the wrong model. it is never deciding, and paying top-tier rates for every job by habit.

stop routing every job by habit

the ai builder toolkit is a set of claude skills that encode how the work gets done in a one-person business: which jobs to delegate, how to scope them so a mid-tier model nails them, and how to check the output before it ships.

see the toolkit

newer to this and want the version with people in the room? the ai builders lounge is where the weekly builds happen.

or just follow along. new field notes most weeks on x, instagram, and tiktok.

one email when the next field note drops.

no course pitch, no daily emails. the notes, when they exist.

written by elisabeth hitz, certified in anthropic's ai fluency program (framework & foundations, and ai capabilities & limitations), plus claude 101 and claude cowork. primary sources: "Introducing Claude Sonnet 5.5," anthropic, published september 28, 2026 (anthropic.com/claude-sonnet-5-5), for the release date, availability, model id, sonnet 5.5 pricing, the "fewer tokens" cost explanation, the opus vs sonnet positioning, the gdpval-aa description, and the slack quote; "Introducing Claude Opus 5.5," anthropic, published september 22, 2026 (anthropic.com/news/claude-opus-5-5), for opus 5.5 pricing and the 40% figure. the gdpval-aa scores of 1,844 and 1,846: venturebeat, "anthropic launches claude sonnet 5.5," september 28, 2026. the models, benchmarks, and pricing are anthropic's; the task-shape routing rule and the four-step test are mine, from setting claude up on real businesses.