Claude Sonnet 4.6 achieves new benchmarks /// OpenClaw hits 50K stars /// Vibe Coding surges 340% /// Claude Sonnet 4.6 achieves new benchmarks /// OpenClaw hits 50K stars /// Vibe Coding surges 340% ///
AI Tools

Shortcut scored 5.9, Claude scored 5.5, both failed

Shortcut scored 5.9, Claude scored 5.5, both failed
A reader asked which spreadsheet agent to buy. The honest answer is a 0.4-point gap on a 10-point test where both tools landed under 60%, plus a $996 annual price difference nobody benchmarked.

Someone dropped the question into Slack this morning with no preamble: Shortcut AI or Claude, for spreadsheet work? Which one. Fair question. The honest answer isn’t a recommendation, it’s a benchmark with a hole in it.

On Wall Street Prep’s 2026 blind evaluation of financial models, Shortcut scored 5.9 out of 10 and Claude scored 5.5. Those two numbers anchor every comparison of these tools I can find. I reached them through a performance comparison of AI financial modeling tools rather than from Wall Street Prep’s own writeup, which I could not locate. No published test set, no count of models graded, no rubric. Two scores, cited in April 2026, and that’s the whole of the independent evidence.

Bon. The gap everyone quotes is 0.4 points. The gap nobody quotes is 4.1, the distance between the winner and a clean model. Hand me a three-statement model I’d grade 5.9 out of 10 and I’m not writing “ship it” at the top. I’m writing “the balance sheet doesn’t close, see me.”

So the interesting question stops being which tool wins. It becomes why both tools score like a first-year analyst on a Friday afternoon.

Shortcut sells itself as the world’s most accurate Excel AI agent and publishes a blind study putting it at 89.1% parity with top-tier analysts. Its side-by-side writeup against Claude puts Claude at 60–70% accuracy on individual formulas, with no way to test its own output. Read that sentence again. The number describing Claude’s accuracy was published by Claude’s competitor. It might be right. It isn’t evidence. Anthropic has published no financial-modeling accuracy figure at all, which is its own sort of answer.

Where 90 percent formula accuracy goes to die

The other Shortcut claim is 90%+ formula accuracy, produced by multi-agent verification. Several passes that check each other’s work. Take it at face value for a second. A three-statement model isn’t a pile of independent formulas. It’s a chain. Revenue feeds gross profit, gross profit feeds EBITDA, EBITDA feeds the cash flow statement, and the cash flow statement is what makes the balance sheet close.

Now the arithmetic. Thirty linked formulas at 90% each, treated as independent, means raising 0.9 to the thirtieth power: about a 4% chance the whole model is clean. Push per-formula accuracy to 99% and the same chain gets you to 74%. Independence is a bad assumption here, since one broken link cascades and the verification passes catch some errors before they land, so treat that 4% as illustration and not measurement. The direction is the point. Per-formula accuracy compounds brutally, and a mid-5s score on the finished workbook is roughly what a high per-formula rate looks like after you multiply it by itself thirty times.

That’s the mechanism behind both scores. These tools are good at the unit and mediocre at the assembly. The failures I’d expect from that pattern are boring and specific: a sign flip on a cash outflow, a hardcoded number sitting in the middle of a formula row, interest expense wired into a circular reference that resolves to zero while nobody notices. It’s not because a model writes a correct XLOOKUP that it can build something a controller will sign.

What the extra $996 a year actually buys

Pro access to Shortcut runs $100 a month billed annually, so $1,200 a year. The free tier hands out 20 weekly credits, and a single message costs 2 to 15 credits, which works out to roughly two real tasks a week before you hit the wall. Claude Pro is $17 a month on the annual plan, $204 a year, with Max between $100 and $200 a month. Line those up. $1,200 against $204 is nearly six times the price for a 7% relative edge, 0.4 points on a base of 5.5, measured once, by someone who didn’t publish a sample size.

Stated that way it sounds damning, and it shouldn’t. Dollars per benchmark point is a stupid metric, and I’d be embarrassed to defend a purchase made on it. What the money buys is location. The plugin runs inside the workbook itself. Claude’s route into Excel is a browser extension listed in the Microsoft marketplace or plain file upload, and while Claude has created and edited xlsx files since October 2025, XLSX upload still sits behind a feature flag according to Anthropic’s own help documentation. An analyst billing at hedge-fund rates who saves one hour a month has covered the spread already. The real case for Shortcut sits right there, and it has nothing to do with 5.9 versus 5.5.

One more framing worth resisting. The tidy conclusion floating around, that these are complementary tools where you plan with Claude and execute with Shortcut, comes from Shortcut’s own comparison post. Vendors love complementary. It moves the conversation off price and onto workflow, and it lets a $1,200 product sit beside a $204 product without anyone doing the division. The framing may well be correct. I’d rather hear it from someone who doesn’t get paid when you agree.

The deployment numbers deserve the same treatment as the accuracy ones. The Shortcut homepage claims three of the five largest multi-strategy hedge funds as customers, more than $100 billion between them. Fundamental Research Labs, the MIT spinout that builds it, raised a $33 million Series A in August 2025, $44.5 million total. Plausible, both. Neither tells you how many seats are active, how many models actually shipped, or how many of those three funds are still running a pilot. Same shape as ServiceNow reporting 9x growth in production agentic-AI customers with no baseline attached. A numerator, loudly, and the denominator left in the drawer.

Worth keeping the base rate in view too. MIT NANDA’s 2025 study of more than 300 enterprise deployments found roughly 95% of generative-AI pilots delivered no measurable P&L impact. Tool choice is not the variable that decides which side of that line you end up on.

Here’s the strongest case against everything I just wrote. A rubric score is not an error rate. Grade a model 5.9 by hand and you might mean the structure was right and three formulas were wrong, fifteen minutes of cleanup rather than a rebuild. I’ve been treating a judgment call as though it were a percentage. That’s a stretch, and I know it. The measurement is stale on top of that: one evaluation from April, on a product from a lab that closed its Series A eight months earlier, tells you little about what either tool does in August. Verification loops improve in weeks. Me, I’ve been wrong about specialist tools before, and I’ve never opened Shortcut, so weigh all of this accordingly.

Don’t buy either one on the benchmark. Take a model you know cold, one where every number is already in your head, and give both tools the identical task. Then time the part nobody times: verification. Not how fast the output arrives, how long it takes you to believe it. Franchement, if checking the work costs more than building it would have, you haven’t bought a modelling agent. You’ve bought a very fast way to be wrong.

Dominic Plouffe

Staff writer at Neural Pulse.