*Most AI news is about the biggest model. Most supply chain work is the same report, every month, done a little faster than last time.*
Anthropic released Claude Sonnet 5.5 today, six days after Opus 5.5. On paper it’s the smaller sibling. In Anthropic’s own tests, it lands within a whisker of Opus on everyday office work, at half the price per token.
That matters more than another record at the top. Nobody runs their carrier tender strategy through an AI every morning. But a lot of us build the same monthly review, supplier pack or KPI deck every few weeks, and that’s where a fast, cheap and careful model pays off.
So this post is about the unglamorous middle of the week. What Sonnet 5.5 is, where it fits next to Opus, and a prompt you can run on a warehouse review today.
What Claude Sonnet 5.5 is
Sonnet 5.5 is the second model in Anthropic’s Claude 5.5 family, released on 28 September 2026. Anthropic positions it as the faster, lower-cost partner to Opus 5.5, which stays the choice for open-ended work that needs sustained judgement. A new Haiku, the smallest model, is promised “in the coming weeks”.
What’s new compared with Sonnet 5, according to Anthropic:
- Faster. Output comes more than 30% faster.
- Cheaper per task. Token prices stay the same, but it uses fewer tokens and fewer tool calls, so a task costs up to 30% less.
- Close to Opus on office work. On GDPval-AA, a test of real tasks from 44 occupations, it scores 1844 Elo, against 1846 for Opus 5.5 and 1449 for Sonnet 5.
- Much better at using a computer. It scores 80.1% on OSWorld 2.1, where the model operates desktop software from screenshots. Sonnet 5 scored 57%.
- Better finished files. Anthropic says it’s strongest at well-scoped tasks and at producing polished documents, slides and spreadsheets.
One early customer quote caught my eye. Box, the document platform, says Sonnet 5.5 rechecks data in the source documents and catches errors Sonnet 5 missed. That is exactly the habit you want in a monthly report.
The price is $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5 and half of Opus 5.5. The context window is 1 million tokens. It’s available in the Claude apps, through the API and on AWS, Google Cloud and Microsoft Azure.
All the benchmark and speed figures are self-reported by Anthropic or its customers. Anthropic itself warns that benchmarks are not a proxy for every production workload.
What it means for supply chain
With two strong models in one family, the real decision is which one gets which job. This is how I’d split the work:
| Kind of work | Supply chain example | Model I’d use |
|---|---|---|
| Open question, new every time | Should we add a sixth hub? How do we restructure the tender? | Opus 5.5 |
| Recurring, well-scoped, checkable | The monthly warehouse review, the supplier scorecard pack | Sonnet 5.5 |
| Thousands of small decisions | Classifying every order confirmation or carrier event | A small or specialised model |
The middle row is the biggest one in most teams, and it’s the row that repeats. A 30% saving per task sounds small once. Over twelve months and twenty reports, it’s the difference between a pilot and a habit.
Where I’d try it first:
- The monthly operating review. Data in, checked KPI workbook and one-page memo out.
- The S&OP pre-read. Turning the demand and supply review into a ten-slide pack people can read before the meeting.
- Inherited planning workbooks. Finding the broken formula in the spreadsheet a colleague left behind, and explaining it.
- Portals without an API. With an 80% OSWorld score, copying booking or ASN data from a carrier portal by screen becomes realistic to test, with a person checking.
- Customer service replies. Drafting answers to order status questions from the order data, fast enough to keep up with the inbox.
The monthly review is the one where carefulness shows. Here’s a test.
Try this: a warehouse review with two traps in it
Upshift is a fictional bicycle maker with two plants and five distribution hubs. Below is two months of warehouse data from its warehouse system, with illustrative numbers. I hid two errors in it, the kind that really happen.
Paste this into Claude with Sonnet 5.5 selected. It needs file creation and code execution switched on, because it asks for an Excel file.
You are a warehouse performance analyst at Upshift, a bicycle maker with five distribution hubs. Prepare the September 2026 monthly warehouse review.
Tasks:
1. Check the data before you use it. Compare the TOTAL row with the sum of the hubs, and look for values that are out of line. Do not correct anything silently: explain each issue and what you did about it.
2. For each hub and month, calculate lines picked per labour hour, pick errors per 1,000 lines and on-time dispatch against the 97% target. Show the change from August to September.
3. Build an Excel workbook with three sheets: the raw data, a KPI sheet that uses formulas (not pasted values), and one clear chart of lines per labour hour by hub for both months.
4. Write a one-page review for the logistics manager: three things that went well, three concerns, and three questions to ask the hub managers.
Rules:
- Use only the data below.
- List every assumption at the end.
DATA
hub,month,orders_shipped,lines_picked,labour_hours,pick_errors,on_time_dispatch_pct
North,2026-08,4820,19300,1610,58,97.8
West,2026-08,3150,12900,1120,41,96.5
South,2026-08,5400,22100,1790,66,98.1
East,2026-08,2280,9600,890,38,94.2
Central,2026-08,6900,28400,2300,85,97.0
North,2026-09,5010,20050,1640,61,98.0
West,2026-09,3080,12700,1105,39,96.9
South,2026-09,5650,23400,1820,70,97.6
East,2026-09,2410,10200,9450,102,89.4
Central,2026-09,7200,29900,2390,90,97.3
TOTAL,2026-09,23530,96250,7900,362,What you should see
- The orders total is wrong. The hubs add up to 23,350 orders, not 23,530. It looks like two swapped digits.
- East’s labour hours are off by a factor of ten. As entered, 9,450 hours gives East 1.1 lines per hour, against about 11 to 13 elsewhere. The TOTAL row of 7,900 hours only adds up if East worked 945, so a good answer uses 945 and says so.
- East is the real concern anyway. Pick errors rose from 4.0 to 10.0 per 1,000 lines, and on-time dispatch fell from 94.2% to 89.4%, far below target. Productivity held at about 10.8 lines per hour on 6% more volume.
- West sits just under target at 96.9%, although it improved from 96.5%.
- The rest looks healthy. North, South and Central pick about 12.2 to 12.9 lines per hour, with around 3 errors per 1,000 lines.
- A workbook with live formulas, so the logistics manager can change a figure and see the KPIs move.
Why this shows the model’s power
A careless model takes the data at face value. It reports that East’s productivity collapsed by 90%, and it copies the orders total into the memo. Both mistakes would send the logistics manager to the wrong hub with the wrong question.
Here the model has to distrust its own input, use one figure to repair another, and then still spot the real problem behind the fake one. Then it has to package it all as a file someone can use. The analyst’s hour goes into asking East what changed in September.
What it can’t do, and the traps
- It can repair a typo, not an invented number. If the warehouse system itself counts labour hours wrongly, the workbook will be neat and wrong. Check one hub by hand every month.
- A near tie on a benchmark is not a tie on your work. Run your last three monthly reviews through both Sonnet and Opus, compare them with what your team produced, and decide per report.
- Well-scoped means you scope it. The prompt above works because the checks, the KPIs and the format are written down. A vague request gets a vague review.
- Fast models make fast mistakes. A computer-use score of 80% still means one task in five goes wrong somewhere. Keep a person on anything that books, orders or pays.
- Mind your data and your settings. Check which account you use and its retention terms before you paste real volumes or customer names. Anthropic offers zero data retention for business use.
Try it on this month’s numbers
This week, take the monthly report your team builds most often and write down its checks, KPIs and format as a prompt. Then run it on last month’s data, next to the version you already sent.
The biggest models get the headlines. For most of us, the real change will come from the middle model that makes every recurring report faster, cheaper and a little more careful.
So tell me: which monthly report would you hand to Sonnet 5.5 first?



