Analytics & Science
Turning that data into decisions.
A single accuracy figure is an average over a test set, and an average is the one summary guaranteed to hide what you need to see. Error concentrates: in rare document types, in the languages with the least training data, and in exactly the ambiguous cases a person would have escalated.
What we build
Ground-truth datasets, and the evaluation sets that tell you whether a system is fit to deploy. Both are people work as much as engineering, which is why the people doing it are permanent employees rather than a task queue.
Measuring the people, not the batch
Every annotator carries a rolling agreement rate, broken down by category. That is what distinguishes one person drifting on one category from a guideline that is genuinely ambiguous — two problems that need opposite responses and look identical in a batch score.
The most useful output is not the ranking. It is the set of items where experienced annotators disagree with each other: those are where the guideline is underdetermined, and they are almost exactly what a model will get wrong later.
What this covers
- Ground truth
- Reference datasets produced by permanent annotators and domain reviewers, with quality measured per person rather than per batch.
- Evaluation design
- Slices chosen because a decision depends on each one, drawn from real traffic rather than curated demonstrations.
- Applied research
- Method work where the off-the-shelf answer does not fit the domain, written up so it can be reviewed.
- Production monitoring
- Drift, cost and refusal behaviour tracked in production, routed back into ground truth.