Measurements
Original measurements from production systems. Every entry is dated, states its method, and says what would change the conclusion. Numbers here; the arguments are in the guides.
51.5%
The Gate Costs More Than the Writing
2026-09-04 · 4 runs · 158 API calls · $22.09
Checking the work cost more than producing it, on every run in the sample. Pure judging was 23.7%; gate-caused revision another 27.8%. One model took 14% of the calls and 51% of the cost.
58.8%
Which AI Writing Tells Actually Show Up
2026-09-04 · 24 gate cycles · 8 raw drafts
Repeated sentence openers were most of everything a mechanical checker caught on raw model output, and the only category that survived revision. The long tail people warn about never fired.
0/10
What Ten Amazon Ads MCP Servers Do With Your Write Access
2026-09-04 · 20 repos searched · 10 audited
None records what it changed after a write. Twenty repositories resolve to far fewer codebases: two are 99.5% byte-identical and GitHub marks neither as a fork.
4/6
Which MCP SDK Version Is Your Client Actually Running?
2026-09-04 · 110 repos resolved · 10 clients audited
Four of the six pinned SDK versions have no tag in the SDK repository, so you cannot read the source your client depends on. One project pins two different versions in the same repo.
An entry belongs here only if it reports first-hand data, states its method well enough to be challenged, carries its collection and revision dates, and says what would change the conclusion. Anything failing one of those is an argument, and arguments belong in the guides.