The Leaderboard Is Not Your Evaluation
A customer asked why we were not using the model at the top of a ranking. We ran the swap on our own set and the higher-ranked model was worse at the work we do: it wrote better prose and followed the schema less reliably, which is exactly the trade a preference ranking rewards. What a leaderboard measures, what it structurally cannot, and how eighty of our own cases decide instead.
