Reference Library
Guides
Durable, maintained references on building agent systems that behave in production. Updated as the ground shifts.
The Agent Reliability Handbook
What it takes to run an agent in production without it quietly breaking things. Contracts, blast radius, replay, and the checklist before you give it write access.
What Horse Racing Taught Me About Model Confidence
Being right is not the goal. Knowing how often you are right is. Calibration, edge, staking, and the discipline of the pass.
Why LLMs Are Bad at Big Data
Language models produce plausible text, not correct numbers. Why bigger context windows do not fix it, and the architecture that does.