Beyond Correctness
AI evaluation is not only about identifying incorrect answers. Every score becomes a signal about the behaviours we want intelligent systems to repeat.
Accuracy has become one of the defining goals of modern artificial intelligence, and for good reason. A model that invents facts, misunderstands instructions or confidently produces the wrong answer will only earn trust for so long. Yet the more time I spend evaluating AI systems, the less convinced I become that correctness is the quality that separates an average response from an exceptional one. Correctness is easier to measure, easier to score and easier to discuss, but usefulness lives somewhere beyond it, and that is where evaluation becomes genuinely interesting.
I have seen responses that answered the question perfectly while completely missing what the person was actually trying to achieve. I have also seen responses that acknowledged uncertainty, made sensible trade-offs and helped someone move forward even though they were not technically perfect. Both may satisfy a rubric, yet only one leaves the interaction better than it found it. That difference is difficult to capture because it is not simply about facts. It sits somewhere between comprehension, context, clarity and judgement, and the more I encounter it, the more I realise that evaluation is not really about deciding whether a response is correct. It is about deciding which behaviours deserve to exist again.
That shift changes the purpose of evaluation. Every score becomes more than a judgement on a single answer; it becomes a signal about the kind of intelligence we are encouraging. Reward a response because it sounds confident rather than because it reasons carefully, and confidence becomes the strategy. Reward length because it appears comprehensive, and verbosity becomes the strategy. Mistake persuasion for understanding, and persuasion becomes the strategy. Models do not decide what kind of systems they want to become. They become increasingly effective at whatever we repeatedly reward, which means every evaluation is shaping their behaviour long before the accumulated effect becomes visible.
The same pattern extends beyond artificial intelligence. People have always responded to incentives in similar ways. Organisations optimise around the metrics they choose to measure, teams adapt to the behaviours that receive recognition, and individuals repeat what is rewarded while gradually abandoning what is ignored. AI makes the principle easier to observe because the relationship between feedback and behaviour can develop at remarkable speed. It reminds us that learning, whether human or artificial, has never been only about acquiring information. It has also been about discovering what matters.
Perhaps that is why judgement has become more interesting to me than correctness. Correctness asks whether an answer is right. Judgement asks whether it is the answer that genuinely serves the person receiving it. The first can often be checked against a reference; the second requires context, restraint, empathy and an understanding that two technically acceptable responses may not be equally valuable. That kind of thinking does not always fit neatly inside a checklist, yet it becomes increasingly important as AI systems move beyond retrieving information and begin helping people make decisions, complete work and navigate situations where the correct response depends on more than factual accuracy.
Working in AI evaluation has unexpectedly become an education in human intelligence because every response reveals something about the standards behind it. Models learn from the signals we provide, but those signals are created by people, which means every evaluation also reflects our own values. Long before an AI system learns what “good” looks like, we have to decide what that word means, which behaviours deserve to be repeated and what kind of intelligence we are trying to build. The clearer our judgement becomes, the better chance our systems have of developing alongside it.