Why frontier models need licensed experts, not crowd raters
As models approach professional-grade performance, the people grading them must be professionals too.
For years, AI training data was produced by large, generalist annotation workforces. That worked when the task was drawing boxes around cars. It breaks down when the task is judging whether a model correctly applied Delaware fiduciary-duty law.
The grader ceiling
A model can only be evaluated as well as its evaluator understands the domain. Once models exceed generalist knowledge, generalist raters stop catching errors — and start rewarding confident, wrong answers.
What changes with experts
Licensed practitioners bring the tacit knowledge that never makes it into textbooks: which edge cases matter, which shortcuts are dangerous, and what a competent professional would actually do.
The takeaway
If your model is meant to operate in a regulated domain, your evaluation data should come from people licensed to practice in it.