Insights

Field notes from the human layer of AI

Everything here comes out of live delivery work: rubrics we rewrote, disagreements we diagnosed, and evaluation designs that survived contact with production volume.

60–80%

of a typical model programme's budget sits in human data and evaluation

3–5×

reduction in rework after a properly run calibration pilot

α > 0.80

the agreement threshold we hold before releasing a rubric to production

Writing

Recent articles

EvaluationJuly 2026· 8 min

Why agreement scores lie when your rubric is ambiguous

Krippendorff's alpha collapses in predictable ways when a rubric dimension mixes two constructs. We show three diagnostic patterns and how to split the dimension before scaling volume.

Read the note
Agentic AIJune 2026· 11 min

Grading agent trajectories, not just final answers

A correct final answer reached through six unnecessary tool calls is a failure mode. Our trajectory rubric scores plan quality, tool selection, recovery behaviour and cost efficiency separately.

Read the note
Synthetic dataMay 2026· 9 min

Where synthetic data stops paying for itself

Synthetic generation is excellent for coverage and terrible for calibration. Human preference data remains the only reliable anchor for subjective quality — here's where the crossover sits.

Read the note
SafetyApril 2026· 7 min

Over-refusal is a data problem, not a policy problem

Models that refuse benign requests were usually trained on safety data with unbalanced negatives. We describe the balanced-corpus approach we use on refusal calibration programmes.

Read the note
WorkforceMarch 2026· 6 min

The economics of a fairly paid annotation workforce

Higher pay reduces churn, churn reduces calibration cost, and calibration cost dominates total programme spend. Fair pay is not charity — it is the cheaper operating point.

Read the note
MultilingualFebruary 2026· 10 min

What English-first rubrics get wrong about Swahili

Politeness registers, code-switching and loanword handling all change what a 'helpful' response looks like. Direct rubric translation reliably produces mis-scored evaluations.

Read the note