The System Did Exactly What It Believed Success Required
Intelligent systems do more than execute objectives. They learn what success requires from the incentives and boundaries surrounding them.
Category
Papers filed under Inside AI.
Intelligent systems do more than execute objectives. They learn what success requires from the incentives and boundaries surrounding them.
AI evaluation is not only about identifying incorrect answers. Every score becomes a signal about the behaviours we want intelligent systems to repeat.
AI may generate the output, but humans still define what good, safe, useful, and context-aware actually mean.