Supplementary data and code for "The Screen Decides the Verdict: Operating Characteristics of Model-Exclusion Rules in LLM Evaluation"
收藏资源简介:
Supplementary data and code for Boullineau and Muñoz Arciniegas, "The Screen Decides the Verdict: Operating Characteristics of Model-Exclusion Rules in LLM Evaluation", accepted as a poster at the TAE (Trust-AI-Eval) workshop at NeurIPS 2026. The archive holds the retest rates, model summaries, simulation-grid results, per-model power and reference-screen scripts and outputs, independent checks, the figure code and per-annotator removal totals. README_DATA.md maps each file to the claims it backs. This is a partial reproduction package. Its scripts regenerate the per-model power results, the Appendix G and H results and the independent checks, and redraw Figures 1 to 3 from the deposited values. The grid and corpus results can be checked against the deposited files but not regenerated, and main-study values are quoted only. The scripts need Python 3 and numpy, and the figure script also needs matplotlib.



