Series 01 / Building better LLMs
The publishing map
Begin with measurement. End with the work that remains uniquely human.
-
01Read ↗
Evals and RL Environments Are the Same Machine With Different Jobs
Prompt, environment, grader: the architecture is shared. The optimization loop is not.
-
02Read ↗
Useful Tasks Live at the Edge of Capability
Too easy produces no signal. Too hard produces only zeros. Progress lives in the narrow band between them.
-
03Read ↗
A Grader Is Never Just Measuring the Model
Every scoring rule creates pressure. The model learns the proxy—especially when the proxy is easier than the goal.
-
04Read ↗
Real Users Have Higher Entropy Than Simulated Ones
Why model-generated users repeat familiar paths—and miss the strange actions that expose real failures.
-
05Read ↗
Long-Horizon Agents Need More Than More Context
A larger window delays forgetting. It does not create judgment, durable memory, or a model of the user.
-
06Read ↗
Continual Learning Is the Missing Layer
Files full of notes are not experience. What would it take for an agent to genuinely get better at its workplace?
-
07Read ↗
Taste Is Becoming the Bottleneck in Model Evaluation
The easy capabilities have abundant data. What remains is difficult precisely because correctness requires judgment.
-
08Read ↗
AI Will Change Software Engineering Before It Replaces Engineers
Automation moves the bottleneck upward—from writing syntax to designing systems, modeling users, and deciding what matters.
-
09Read ↗
What a Good Agent Trajectory Must Capture
The state, actions, evidence, provenance, and costs required to explain why an outcome happened.