01
"You upgraded your model. What broke?"
Statistically rigorous drift detection between LLM endpoints. It scores eval suites with paired significance tests, prices every verdict in cost per correct answer, and runs a scheduled observatory that watches live endpoints for silent model swaps.
I use it to document drift between frontier releases from the outside. One finding: a silent tokenizer change raised a model's cost per correct answer 35% while accuracy barely moved.