JOURNAL

Claude Sonnet 5.5 vs Opus 5.5: I tested both on the same sales forecast

Same messy CRM export, same prompt, nine kinds of planted error. Both hit the number to the krone. In my run Sonnet was a minute faster, and it's half the price per token. Opus built the better deck.

29 September 2026·9 min read·Claude · Claude Sonnet 5.5 · Claude Opus 5.5 · RevOps · AI for revenue operations

TL;DR: I tested how good Claude Sonnet 5.5 is on a realistic piece of sales work. It got a messy CRM export with 61 deals and nine kinds of error I'd planted on purpose, and had to deliver a Q4 forecast in Excel and five slides for the board. Then Opus 5.5 got exactly the same job. Both hit the forecast to the krone, DKK 6,529,724, and both caught all nine error types. On my clock Sonnet finished in 2 minutes 3 seconds, Opus in 3 minutes 16, and at API list prices the runs cost $0.58 against $1.39. Opus built the better deck. My takeaway: Sonnet for the analysis you run every week. Opus for what the board sees.

Both runs, sped up. The clock shows real time.Brinvik

Claude Sonnet 5.5 came out on 28 September. I've written about what the switch to Sonnet 5.5 means for your agents. But what I wanted to know was something else: how good is it at the work a sales team actually does?

So at two in the morning I put it to work as a RevOps analyst. Then I gave exactly the same job to Opus 5.5, Anthropic's pricier model for harder work. Same prompt and same file, sent in two Claude app windows at the same moment. The film above shows the two runs side by side, rebuilt from the run logs and the real files and sped up. The clock is real time.

The job: a messy CRM export and a board meeting

The company is called Havblik Software ApS. It doesn't exist. Neither does the pipeline, but it looks like one I've seen many times: 61 deals exported from a CRM on a Tuesday morning because the CFO needs a number for the board.

The model had to do three things. Clean the data without silently dropping anything. Build a weighted Q4 2026 revenue forecast in Danish kroner with fixed probabilities per stage, from 10 percent for Qualified to 70 percent for Negotiation, and euros converted at 7.46. And deliver three files: an Excel workbook with live formulas, five slides a CFO can present as is, and a list of every data problem.

A spreadsheet with rows from the real export. The errors are circled in red: an amount of 7,200,000 with the note customer budget approx 720k, the Danish stage name Tilbud sendt, the amount 765.000,00 stored as text, €46,000 with an empty currency column, a blank close date, a probability of 95% on a Negotiation deal, an open deal with a close date of 30 June, a duplicate row from an old CRM and a Closed Won deal dated 8 November 2026.
A slice of the export. Nine kinds of error, planted on purpose, and I've seen every one of them in a real CRM.

There were nine kinds. Three deals appeared twice after an old import. Four amounts were in euros, and on two of them the currency was only a € sign. Stage names came in five non-standard spellings, two of them Danish. Five open deals had a close date that had already passed, and four had none. Three probabilities contradicted the stage. Four amounts were stored as text in Danish number format. One deal had an extra zero. And one was marked won with a close date in November.

Before either model saw the file, I wrote the answer key: DKK 6,529,724 from 27 deals, and every error by deal ID. That way neither of them could talk me round with a nice slide.

Same forecast, to the krone

Both models landed on DKK 6,529,724 from 27 deals. Exactly the answer key. Both caught all nine error types, kept all 61 rows and built the workbook with real formulas that recalculate if you change a probability. Neither of them typed the total in as a number.

They also made the hard calls the same way. The deal at DKK 7.2m with a note saying the budget was about 720k was corrected to 720,000 and flagged for the rep. The duplicates were kept out and the originals kept in. Deals with no close date stayed out of Q4, because you can't put them in a quarter.

Where they differed was in what else they noticed. Opus spotted that two of the duplicates carried Q4 close dates while the originals had dates in the past, and worked out that the forecast rises by DKK 0.30m if sales confirms the new dates. Sonnet put the same 0.30m on its upside slide, and also questioned a deal at 75,000 that might have been 750,000. The 75,000 was actually right. But that's the kind of doubt I'd rather have in writing than buried.

Sonnet 5.5 was faster and half the price per token

Sonnet 5.5 finished in 2 minutes 3 seconds with 9 tool calls. Opus 5.5 took 3 minutes 16 seconds and 15 calls. That's my own measurement from one run each, and timing varies from run to run.

Animation on a light ground. Two lanes, Claude Sonnet 5.5 and Claude Opus 5.5, fill with a mark for each tool call, 9 for Sonnet and 15 for Opus, while a clock counts up. Sonnet stops at 2:03, Opus at 3:16. Both lanes end on the same number: DKK 6,529,724.
Same job, same result. Sonnet 5.5 finished a minute ahead of Opus 5.5 with six fewer tool calls. Time and call counts are my own measurement from one run.

The price doesn't vary. Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens. Opus 5.5 costs $4 and $20. Half the price per token, for the same number.

At Anthropic's list prices, from the token counts in each run's log, the runs cost $0.58 for Sonnet and $1.39 for Opus. Sonnet's output tokens are an estimate. So Opus cost almost two and a half times as much, because it's pricier per token and it used more of them.

I haven't dug into what Opus spent its six extra calls on. What I can see is what it delivered.

Opus 5.5 built the deck I'd show a board

Opus made five slides that hang together. Three headline numbers up top. A bar chart with a stage table beside it. A slide showing that December carries half the quarter, DKK 3.37m of 6.53m, next to the quarter split by sales rep. A slide on the cleanup with three cards saying what it did to the number. And finally risks, upside and what sales has to fix by 6 October.

The two real decks side by side, five slides each. Opus 5.5 has headline numbers in cards, charts with tables, a split by sales rep and a slide on December. Sonnet 5.5 has three big numbers, one bar chart and then three slides of bullet points.
Both decks exactly as the models delivered them, unedited. Same numbers, very different presentation.

Sonnet's deck was correct. The numbers were there, the conclusions were right, and slide 4 on what could move the number was genuinely sharp. But after the first two slides it was bullet points. A CFO could use it. Nobody would look forward to presenting it.

That matches what Anthropic said at launch: Sonnet 5.5 is strongest at well-scoped everyday tasks, while Opus 5.5 is still clearly stronger at complex, open-ended work that needs judgment over time. A board deck is judgment: deciding what goes in and what the board should worry about.

How I'd split the models in AI for revenue operations

Pick the model per deliverable, not per company. That's where I landed after the test, and it isn't where I started. I expected to find a winner.

Instead I found two different jobs. On the weekly analysis, where the number has to be right and the errors have to be found, Sonnet did as well as Opus, faster and at half the price per token. On the thing that goes in front of a board, Opus did better, and there an extra minute and double the token price don't matter.

That's AI for revenue operations in practice, and it's how you reduce repetitive work with AI: the dull part runs every week on the cheaper model, and you save the expensive one for where judgment counts.

A working week, Monday to Friday. Every day the recurring work runs on Claude Sonnet 5.5: clean the pipeline, refresh the forecast, deal hygiene. On Friday, at month end, the board deck is marked in red and runs on Claude Opus 5.5.
Sonnet for what you do every week. Opus for what the board sees.

One thing doesn't change whichever model you pick. Both models flagged deals a human has to answer: the amount with an extra zero, the won deal dated in the future, the three duplicates. A model can find the error. It can't ring the rep and ask.

What the test doesn't show, and how to run it yourself

One run per model, on data I made up, against an answer key I wrote. That makes it a first look. I timed it myself, and another run could give another number. Nothing broke along the way, but Sonnet's doubt about the 75,000 deal was a false alarm, and in a real pipeline a false alarm costs a rep's time.

So you can take the whole test and run it yourself. Download the test kit here. It has the prompt both models got, the CSV with the 61 deals, the answer key and both models' files exactly as delivered. Swap in your own export, anonymised, and write your own answer key before you look at the result.

In my piece on Opus 5.5 I argued that the only measurement that counts is the one you run on your own work. This is that measurement, run once. It won't tell you anything about your data until you've run it on your data.

Internal AI tools for sales teams: what I build

As an AI consultant I build this kind of thing for sales teams: an agent that cleans the pipeline and refreshes the forecast every week, and writes the reps a list of what only they can answer. The hard part is usually writing the answer key, so you can tell whether the agent is right before you let it run on its own. That's what I mean by internal AI tools, and it's where Claude for business pays off most in a sales team.

FAQ

Frequently asked questions

In my test, yes. It hit the weighted Q4 forecast exactly, DKK 6,529,724 from 27 deals, and caught all nine planted error types. That's one run on made-up data, so run the test on your own export before you trust it.

Half as much per token. Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens. Opus 5.5 costs $4 and $20. In my test the tokens used came to $0.58 for Sonnet and $1.39 for Opus at list prices.

When the deliverable needs judgment: what goes in, what gets emphasis, what the reader should worry about. In the test Opus built the better board deck. Anthropic says so too: Opus 5.5 is still clearly stronger at complex, open-ended work.

Yes. The test kit has the prompt, the CSV and the answer key. Use an anonymised export, write your own answer key first, and check your data processing agreement before real customer data goes to an AI model.

Yes, both. The forecast sheet calculates with COUNTIFS and SUMIFS against the clean data sheet, so the number changes if you correct a probability or a deal. Neither typed the total in as a fixed number.

Start with one job you already do every week, such as pipeline cleanup or the forecast. Write an answer key for a week you know the answer to, and let Claude run it. Once the number's right, it can run on a schedule and you spend your time on the deals it flags.

It depends on how many systems the agent has to talk to. I work at a fixed price, and a weekly forecast agent on one CRM is a well-defined job. The prices are on my pricing page.

Sources

How this article was made. This work was produced together with AI. Overall: AI about 71 percent, Kim about 29 percent. Counting production alone, AI did about 92 percent. Kim ran the test, timed it and set the angle. Claude wrote the articles, built the graphics and re-checked the numbers against the test's own files. Kim wrote the rules, the voice and the method in advance. The numbers are a qualified estimate, not a measured log.

Portrait banner on a light ground. The headline "Who did the work?", below it two numbers for AI and Kim, a bar divided in the same proportion, and phase bars with their own labels underneath.
The split is a qualified estimate over the finished task, not a measured log.

Get new essays by email.

Roughly twice a month. Same voice. No list rental, no retargeting.

Protected by Cloudflare Turnstile. No challenge, no CAPTCHA. Brinvik never shares your address.