← Post-training series

Essay 03 / Task design

Useful tasks live at the edge of capability.

Too easy gives us applause. Too hard gives us silence. The useful region is where present systems disagree.

SERIES 02 / ESSAY 03

03
TASK DESIGNTHE USEFUL EDGE

TOO EASY ← SIGNAL → TOO HARD

  • Variance
  • Real failure
  • Moving frontier

A task is useful only when its score can change a decision. That requires enough difficulty to expose a gap, but enough possibility for capable systems to show what they know.

POST-TRAINING MODELSESSAY 03 / 10

Perfect scores and zero scores tell the same story.

When every model passes a task, the result contains no separation. The task may still demonstrate a feature, but it no longer tells a researcher which system is more capable. At the other extreme, universal failure compresses every model into the same zero.

Both are measurement dead zones. The useful region lies between them: a place where strategies, robustness, and judgment cause observable differences. That variation is the raw material of a good decision.

Difficulty is not the goal. Informative variation is.
Too easyNo separation
Useful edgeBehavior varies
Too hardNo foothold

Begin with a failure you can actually feel.

Abstract lists of capabilities are tempting because they look systematic. But the best tasks usually begin more simply: try to use the model for real work and notice where trust breaks. Perhaps it loses track of a long workflow, edits the wrong state, or succeeds only when the instructions remove all ambiguity.

A real failure already contains context about why the capability matters. The work is to reduce it into a reproducible task without sanding away the source of difficulty. If the original problem required judgment across several steps, turning it into a single obvious choice produces a cleaner task and a weaker measurement.

A useful source task has:

  • A concrete user or researcher pain point.
  • An outcome that can be observed consistently.
  • Enough information for a fair attempt.
  • More than one plausible path through the work.

Do not simplify away the capability.

Task construction always removes details. Some details are noise; others carry the challenge. The central design judgment is knowing which is which. If we make navigation trivial, provide the key intermediate fact, or reduce a long sequence to a one-step command, we may accidentally measure instruction following instead of the intended capability.

A strong task keeps the causal structure of the failure intact. The model should encounter the same kind of uncertainty, delayed consequence, or competing objective that made the original situation difficult. Reproducibility should control irrelevant variance, not eliminate meaningful complexity.

  1. 01
    Name the capability.

    Write down the exact behavior the task is meant to expose.

  2. 02
    Identify the load-bearing difficulty.

    Find the detail whose removal makes the task suddenly easy.

  3. 03
    Test alternate solutions.

    Confirm that success reflects the capability, not one memorized path.

A good task is designed to become obsolete.

Capability moves. A task that once separated models may eventually saturate, while a task that once looked impossible may enter the useful band. This is not instability to hide. It is the reason task suites need maintenance.

Teams sometimes protect an old benchmark because its history makes comparison convenient. Historical continuity matters, but it cannot replace present diagnostic value. Keep a small anchor set for trend lines, then invest most attention where contemporary systems still reveal meaningful differences.

The result is a living measurement program: watch real use, collect failures, locate the edge, and retire tasks once their signal is exhausted.

Build for the decision after the score.

The right level of difficulty is not a fixed percentage. It depends on what the result must help someone decide. Useful tasks make capability differences visible, preserve the reason those differences matter, and move as the models move.

View the complete series
Next · Essay 04

A Grader Is Never Just Measuring the Model

How scoring rules become training pressure.