Writing

I publishwhat Imeasure

Essays and benchmarks on agent systems, evaluation, and context. All of it lives on my Substack.

Read on Substack
Fig. 001

What I write about

Series

Minimal Viable Context

Runnable context packages that replace PRDs. What an agent actually needs to do the work, and nothing else.

Read the series

Benchmarks

Rift findings

What breaks when frontier models get upgraded, measured from the outside with my own CLI.

Read the findings

Economics

The unit economics of intelligence

Cost per correct answer, and why quality drift is a budget problem before it is a quality problem.

Read the essay
Fig. 003

Subscribe

One click to confirm on Substack. Essays and benchmarks, nothing else.