遇见数据集

Supplementary data and code for "The Screen Decides the Verdict: Operating Characteristics of Model-Exclusion Rules in LLM Evaluation"

收藏
Zenodo2026-09-24 更新2026-10-01 收录
官方服务:

资源简介:

Supplementary data and code for Boullineau and Muñoz Arciniegas, "The Screen Decides the Verdict: Operating Characteristics of Model-Exclusion Rules in LLM Evaluation", accepted as a poster at the TAE (Trust-AI-Eval) workshop at NeurIPS 2026. The archive holds the retest rates, model summaries, simulation-grid results, per-model power and reference-screen scripts and outputs, independent checks, the figure code and per-annotator removal totals. README_DATA.md maps each file to the claims it backs. This is a partial reproduction package. Its scripts regenerate the per-model power results, the Appendix G and H results and the independent checks, and redraw Figures 1 to 3 from the deposited values. The grid and corpus results can be checked against the deposited files but not regenerated, and main-study values are quoted only. The scripts need Python 3 and numpy, and the figure script also needs matplotlib.

提供机构:
Zenodo
创建时间:
2026-09-24
二维码
社区交流群
二维码
科研交流群
商业服务