Skip to content
Theory AI
← All services

Analytics & Science

Turning that data into decisions.

A single accuracy figure is an average over a test set, and an average is the one summary guaranteed to hide what you need to see. Error concentrates: in rare document types, in the languages with the least training data, and in exactly the ambiguous cases a person would have escalated.

What we build

Ground-truth datasets, and the evaluation sets that tell you whether a system is fit to deploy. Both are people work as much as engineering, which is why the people doing it are permanent employees rather than a task queue.

Measuring the people, not the batch

Every annotator carries a rolling agreement rate, broken down by category. That is what distinguishes one person drifting on one category from a guideline that is genuinely ambiguous — two problems that need opposite responses and look identical in a batch score.

The most useful output is not the ranking. It is the set of items where experienced annotators disagree with each other: those are where the guideline is underdetermined, and they are almost exactly what a model will get wrong later.

What this covers

Ground truth
Reference datasets produced by permanent annotators and domain reviewers, with quality measured per person rather than per batch.
Evaluation design
Slices chosen because a decision depends on each one, drawn from real traffic rather than curated demonstrations.
Applied research
Method work where the off-the-shelf answer does not fit the domain, written up so it can be reviewed.
Production monitoring
Drift, cost and refusal behaviour tracked in production, routed back into ground truth.

Working on something like this?

Tell us what you are trying to deploy and what data you hold.

Bring us the use case