Count the gradings, not the questions: how far each conclusion of a small paired evaluation is from being overturned
收藏官方服务:
资源简介:
Per-answer question-answering records under three retrieval conditions, paired contrasts, an independent three-grader regrade, minimum-flip and leave-one-out summaries, figure code and manuscript source for a fragility audit that computes the smallest number of individual gradings that would overturn each reported direction or verdict. Local KG results are not DocBench accuracy. Code is MIT; manuscript text, figures and recorded result data are CC-BY-4.0. Third-party simulators and datasets are not redistributed.
提供机构:
Zenodo创建时间:
2026-09-10



