633 points on Hacker News, for a model you cannot chat with. That is where the Hacker News thread sat two days after Cloudflare posted it. The model itself answers in milliseconds and gives you back a number, not a sentence — that’s the whole shape of an open-weight decision model. Four days earlier, AWS’s Strands Labs quietly shipped almost the same idea, smaller, and got basically none of that attention. Bon. Let’s talk about what actually shipped, because the category matters more than either announcement.
Me, I’ve written this exact bad pattern more times than I want to admit. You need a yes/no, or a score, or a pick-one-of-five, somewhere in a pipeline. Reorder this SKU or not. Flag this listing for review or not. Approve this price change or not. And the fastest way to get there, when you already have an LLM wired into your stack, is to prompt it: “respond with JSON, one field called decision, one field called confidence.” Then you parse that JSON and pray the model didn’t wrap it in a markdown fence this one time, or add a sentence of preamble before the braces, or return “yes” where your schema expects “true.” It works.
Until it doesn’t, and you find out in production, at 3am, because the one time it mattered is the time the format drifted — the same way a bug I once chased was perfectly reproducible while my logs told me nothing useful about why.
That is the gap Clef and Strands Decider are both built to close, and reading the two model cards back to back, you can see two different vendors arrive at nearly the same architecture independently, which is the actual story.

What Clef’s open-weight decision model actually does
Cloudflare’s post describes two sizes: Clef, built on a Qwen 3.8-27B backbone, and Clef-flash, built on Qwen 3.5-9B. Neither is trained from scratch. Both are a frozen base model plus a rank-256 LoRA adapter, with a scoring head that does something genuinely different from a normal chat model — instead of generating tokens one at a time, it runs a non-autoregressive, two-stage attention pass that scores every allowed answer in your schema in parallel, then picks. The training loss is label-smoothed cross-entropy plus a Brier loss term specifically to make the probability numbers calibrated, not just the top pick correct. That Brier term is the detail that tells you this was built by people who actually want to threshold on the confidence number later, not just read the top label.
The numbers that matter to me: median latency of 209.3ms for Clef, 38.8ms for Clef-flash, with p95 at 238.6ms and 122.4ms. 64k context window on both. Clef-flash scores 98.76% on BFCL and 97.73% on a “Home Appliances” benchmark Cloudflare ran themselves — so take that second one as their own number until someone replicates it. Both are Apache 2.0, weights on Hugging Face, and you can run them yourself off the repo instead of only through Cloudflare’s Workers AI.
Strands Decider 2B is the smaller, less-covered sibling
AWS’s Strands Labs posted the Strands Decider announcement on the same day, October 1st, and it is a more aggressive shrink of the same idea. It starts from Qwen3.5-2B, rips out the language modeling head entirely, and bolts on a pointer head of just over a million parameters, trained with a rank-16 LoRA. Total model: 2 billion parameters. Median latency around 115ms on an RTX 3090, 153ms on an M3 MacBook — hardware a lot of us actually have under a desk, not a data-center card. On JevBench’s public leaderboard it lands 3rd of 33 models in the 2B class, 1st of 30 if you exclude anything just over the 2B line, on both accuracy and the calibration score. The repo is on GitHub, weights are on Hugging Face under the StrandsAgents org, Apache 2.0 licensed, and as of this week it sits at 331 stars — small, early, real.
Nobody on Reddit or Hacker News is arguing about Strands Decider yet, which is itself a data point. Clef got the 633-point thread and the r/LocalLLaMA pickers going through the weights within a day. Strands Decider got a VentureBeat writeup and basically nothing from the communities that actually run this stuff. Same week, same idea, wildly different attention — that tells you more about who posted first and which vendor’s name is bigger than it tells you about which model is better.
Is “frozen backbone plus LoRA adapter” actually open
This is the argument running through the HN thread, and I’ll take a side on it instead of restating both. A chunk of commenters are annoyed that Clef isn’t trained from scratch — it’s Qwen with a 256-rank adapter bolted on, so calling it “Cloudflare’s model” feels like a stretch. I think that complaint misses what matters for someone trying to actually use the thing.
The weights are published, the license is Apache 2.0, you can download it and run it with no API key and no per-token bill to Cloudflare, and the adapter plus scoring head are exactly the part that makes it a decision model instead of a chat model. Whether Cloudflare trained 27 billion parameters or borrowed 27 billion and trained 256 ranks’ worth, I can put it on my own GPU tonight. That is the test for “open” that matters to a startup, not who gets academic credit for the pretraining run.
The question I actually care about
Would I swap one of my LLM-calls-pretending-to-be-a-classifier for one of these. Honestly, maybe, for the narrow cases — flag-for-review, approve-this-price-change, the stuff that’s genuinely closed-set. What would break: these models are fine-tuned for their own benchmark tasks, not mine, so I’d need my own labeled examples before the calibration numbers mean anything for my schema, and that labeling cost is real, not theoretical. A rank-16 or rank-256 adapter on a frozen backbone also means you inherit whatever the base Qwen checkpoint is bad at — there’s no free lunch where the adapter fixes a backbone weakness it was never trained to touch. And at 2 to 27 billion parameters, none of this replaces a model that needs to reason through something novel; it replaces the part of my pipeline that was never reasoning in the first place, just dressed up as if it were.
Three teams converging on frozen-backbone-plus-adapter-plus-calibrated-head inside the same week isn’t a coincidence, it’s a sign the shape is right. The part worth watching isn’t the hero number on a model card, it’s whether the next six months brings a decision model fine-tuned on a task you actually have, with weights you can check yourself instead of a vendor’s chart.