writing
Notes from the build room
Write-ups of work that already shipped, with the measurements attached. Mostly evaluation: how you know an AI product is doing the thing you claim it does.
- ·agent graphs
The model cannot own the money
Dozen is a twelve-node creator marketplace. Accept, post, and pay stay outside the model. This is what the graph actually owns, what Stripe Checkout does, and what is still scaffolding.
read it → - ·llm evaluation
Benchmarking a generative character when there is nothing to diff against
A shipped AI character, an 82% hidden-reasoning bill, and a whole-game evaluation harness for the quality a one-shot test cannot see.
read it → - ·ai crawlers
Your firewall is your AI policy
I probed 18 major sites with the user-agent of every AI crawler that matters. Who gets a 200 and who gets a 403 lines up with who signed deals and who is in litigation — and two of my own findings did not survive the data.
read it → - ·llm evaluation
Calibrating an LLM judge for a game people are trying to beat
A scoring model that grades live players has no reference implementation to check itself against. What I measured, what broke, and the calibration metric I had to change.
read it →