writing
Notes from the build room
Write-ups of work that already shipped, with the measurements attached. Mostly evaluation: how you know an AI product is doing the thing you claim it does.
- ·ai crawlers
Your firewall is your AI policy
I probed 18 major sites with the user-agent of every AI crawler that matters. Who gets a 200 and who gets a 403 lines up with who signed deals and who is in litigation — and two of my own findings did not survive the data.
read it → - ·llm evaluation
Calibrating an LLM judge for a game people are trying to beat
A scoring model that grades live players has no reference implementation to check itself against. What I measured, what broke, and the calibration metric I had to change.
read it →