We propose confidence as a first-class scaling objective. Traditional evaluation asks whether a model is right. Our research asks whether it remains composed when it is not.
In this conceptual framework, additional parameters expand the space of plausible explanations without imposing the inconvenience of additional evidence. A wrong answer can therefore become progressively more articulate.
The central open problem is not confidence itself, but its unfortunate correlation with credibility. We recommend evaluating assertions against evidence rather than evaluating evidence against the eloquence of assertions.