Easy-to-measure capabilities are often the first to saturate. What remains is not merely harder for models; it is harder for us to describe, judge, and convert into dependable tasks.
The remaining failures are increasingly about judgment.
Models can execute many well-specified transformations with impressive consistency. The difficult cases now cluster around incomplete intent: what should be clarified, which constraint is truly important, when a technically correct answer would still be unhelpful, and when the safest action is to pause.
These capabilities do not yield easily to a checklist. Two experts may agree that one response is better while struggling to express a rule that generalizes to the next situation. That tacit standard is taste: learned judgment about what good work looks like in context.
When the specification is incomplete, the evaluator becomes part of the specification.
High-context task creators are hard to scale.
A person who has repeatedly done the work can recognize subtle failure modes. They know which shortcuts create downstream pain and which imperfections are harmless. Turning that knowledge into a task requires time: isolate the scenario, make it fair, build the environment, write scoring logic, and inspect model behavior.
The feedback loop is long, and the people with the most relevant judgment are often the busiest. Scaling by using less experienced creators can increase volume while flattening the very distinctions the task was meant to preserve.
Task quality needs:
- Experience with the real work and its downstream consequences.
- A clear account of why one outcome is better than another.
- Examples that challenge the creator's first scoring rule.
- Review by people with different but relevant perspectives.
More judgments do not erase systematic taste.
If individual ratings contain random error, aggregation helps. But evaluators may share the same blind spot: a preference for polished prose over correct restraint, for conventional solutions over effective unusual ones, or for visible activity over quiet prevention.
That bias becomes more dangerous as scoring moves closer to subjective quality. Agreement can reflect a shared culture rather than universal usefulness. Diverse reviewers help, but diversity must map to the users and consequences the system will face, not exist as a decorative statistic.
Rationales are valuable here. Asking why a judgment was made exposes hidden assumptions and makes disagreement diagnostic. The goal is not to eliminate taste; it is to make its sources visible enough to audit.
Checking hard work can require the same capability as doing it.
Verification is easy when success has a crisp external state. It is much harder when the task asks for a good product decision, a sensitive escalation, or a maintainable design. A model capable of producing plausible work may also be capable of producing plausible explanations for weak work.
This symmetry limits fully automated evaluation. Stronger judges help, but they do not remove the need for task design, counterexamples, and periodic human review of downstream behavior.
- 01Anchor in real pain.
Start from consequences users have experienced, not abstract preference.
- 02Collect disagreement.
Use contested cases to discover missing dimensions of quality.
- 03Demand rationales.
Make the assumptions behind judgments inspectable.
- 04Audit after deployment.
Compare the score with what happens in actual use.
Better models require better standards of good.
As obvious capabilities improve, evaluation becomes less like counting correct answers and more like articulating mature professional judgment. The bottleneck is not only model intelligence. It is our ability to recognize, explain, and consistently reward the behavior worth building.
View the complete series