Per-judge score matrices and generated answers for a multi-judge study of small-LLM response aggregation (8 paradigms x 400 queries x 3 judges)
收藏资源简介:
Companion artifacts for the paper "Robust at the Top, Fragile in the Middle: A Multi-Judge Study of Small-LLM Response Aggregation Rankings". Eight response-aggregation paradigms were re-implemented behind one abstract interface and run over an identical expert pool and a 400-query multilingual test set (Korean, English, German, Japanese; 100 queries each). The resulting answer set was then frozen and rescored by three independent LLM judges (Claude Haiku 4.5, Sonnet 4.6, Opus 4.7), so that any change in the reported ranking is attributable to the judge alone. This record contains the complete per-query, per-judge score matrix (8 x 400 x 3 = 9,600 scores), the generated answers with their measured latency and cost, the derived agreement and significance statistics (Cohen's kappa with query-clustered bootstrap confidence intervals, Spearman and Kendall rank correlations with exact permutation p-values, Holm-adjusted paired Wilcoxon verdicts), the five figures of the paper, and the analysis code that produces all of them. Source questions are not included. The test set draws on four public corpora whose redistribution terms differ, and the Korean partition in particular permits release of item identifiers and generated output only. Every item identifier is published, so any item can be retrieved from its upstream corpus. Runs of 60 or more characters copied verbatim from a source question into an answer are masked; the counts are recorded in manifest.json.



