Quality isn't reviewed at the end. It's engineered in.
Most data programmes fail on ambiguous rubrics, not lazy annotators. Our process front-loads guideline design and calibration so production volume is trustworthy from week one.
Six phases from brief to steady state
Scoping & feasibility
We map your task type, difficulty distribution, expected volume and acceptance criteria, then model realistic throughput and the tasker profile required.
Output · Delivery plan + capacity model
Rubric & guideline design
Our taxonomists co-author annotation guidelines with your researchers, resolving edge cases before they become inter-annotator disagreement.
Output · Versioned guideline doc
Pilot & calibration
A 500–2,000 task pilot measures inter-annotator agreement (Krippendorff's alpha), throughput and rubric ambiguity hotspots. Guidelines are revised before scale.
Output · Agreement report + revisions
Team assembly & training
Taskers are screened, assessment-tested on your rubric, then certified. Only certified taskers receive production access on the client platform.
Output · Certified production roster
Production delivery
Live queues run with gold-task injection, blind overlap sampling, reviewer tiers and daily throughput dashboards visible to your team.
Output · Daily throughput + QA metrics
Audit & iteration
Weekly quality councils review disagreement clusters, retrain where drift appears and feed rubric updates back into training material.
Output · Weekly quality council notes
Four independent layers of verification
Each layer catches a different failure mode: inattention, ambiguity, systematic bias, and mismatch with the client's actual intent.
Layer 1 — Gold tasks
Known-answer items seeded at 4–7% of volume, scored automatically. A tasker falling below threshold is paused and re-trained, not silently degraded.
Layer 2 — Blind overlap
10–20% of tasks are double- or triple-assigned without the taskers knowing, producing live inter-annotator agreement statistics per rubric section.
Layer 3 — Reviewer tier
Senior reviewers adjudicate disagreements, sample high-risk categories at 100%, and write the disagreement notes that drive guideline updates.
Layer 4 — Client acceptance
A final sample is packaged for your researchers with rationale attached, so acceptance decisions are made on evidence rather than aggregate scores.
What lands in your inbox every Monday
Inter-annotator agreement
Krippendorff's α, per rubric dimension
Gold-task accuracy
Rolling 7-day per tasker and per pod
Throughput
Accepted tasks/hour vs. capacity model
Rework rate
% returned by reviewer or client
Latency
Median and p95 queue-to-delivery time
Drift
Week-over-week rubric-section score movement