← Back home

Evals · RL environments · agents

Ideas,
in sequence.

9 essays · complete series

Nine compact field notes on measuring models, designing environments, and building capable agents.

See every title

Series 01 / Building better LLMs

The publishing map

Begin with measurement. End with the work that remains uniquely human.

  1. 01
    Published · 13 Aug 20265 min read

    Evals and RL Environments Are the Same Machine With Different Jobs

    Prompt, environment, grader: the architecture is shared. The optimization loop is not.

    Read
  2. 02
    Published · 13 Aug 20265 min read

    Useful Tasks Live at the Edge of Capability

    Too easy produces no signal. Too hard produces only zeros. Progress lives in the narrow band between them.

    Read
  3. 03
    Published · 13 Aug 20265 min read

    A Grader Is Never Just Measuring the Model

    Every scoring rule creates pressure. The model learns the proxy—especially when the proxy is easier than the goal.

    Read
  4. 04
    Published · 13 Aug 20265 min read

    Real Users Have Higher Entropy Than Simulated Ones

    Why model-generated users repeat familiar paths—and miss the strange actions that expose real failures.

    Read
  5. 05
    Published · 13 Aug 20265 min read

    Long-Horizon Agents Need More Than More Context

    A larger window delays forgetting. It does not create judgment, durable memory, or a model of the user.

    Read
  6. 06
    Published · 13 Aug 20265 min read

    Continual Learning Is the Missing Layer

    Files full of notes are not experience. What would it take for an agent to genuinely get better at its workplace?

    Read
  7. 07
    Published · 13 Aug 20265 min read

    Taste Is Becoming the Bottleneck in Model Evaluation

    The easy capabilities have abundant data. What remains is difficult precisely because correctness requires judgment.

    Read
  8. 08
    Published · 13 Aug 20265 min read

    AI Will Change Software Engineering Before It Replaces Engineers

    Automation moves the bottleneck upward—from writing syntax to designing systems, modeling users, and deciding what matters.

    Read
  9. 09
    Published · 13 Aug 20266 min read

    What a Good Agent Trajectory Must Capture

    The state, actions, evidence, provenance, and costs required to explain why an outcome happened.

    Read