The Epistemic Benchmark is a fictional evaluation design for the gap between evidence quality and expressed certainty. It distinguishes missing information from information a model would prefer not to acknowledge.

Example tasks include revising an assumption, admitting that a comparison is unsupported, and resisting a consensus that has no identifiable source. None of the scores on this site represent measured model performance.

Our proposed reporting standard places uncertainty beside every conclusion. Early conceptual work suggests that readers may find this substantially less exciting.