遇见数据集

Coded corpus of 86 paired comparisons between published LLM claims and cross-model retests, with four-model judgment experiments

收藏
Zenodo2026-09-21 更新2026-10-01 收录
官方服务:

资源简介:

This dataset supports the manuscript "When should new results change a published claim?". The study assembled 75 original research papers and 34 distinct follow-up papers (109 papers in total) that retest published comparative claims about large language models after model updates, formed 86 paired cases, coded each pair against seven comparison conditions, and assigned each case one of four judgments. Two annotators coded all cases independently, disagreements were adjudicated, and four large language models were tested on the judgment task under several input conditions. The deposit contains the coded corpus with per-case condition tables and an index of the underlying papers, the coding rules, the annotation and adjudication materials, a screening summary, the per-case model predictions and scores from which all reported metrics can be recomputed, the attribution analysis, and the core scripts. Full texts of the collected papers are not redistributed due to copyright and can be retrieved through the identifiers in the case index.

提供机构:
Zenodo
创建时间:
2026-09-21
二维码
社区交流群
二维码
科研交流群
商业服务