遇见数据集

Auditing the Auditor: A Set-Overlap Protocol for Validating Quality Proxies in AI-Generated Information Pipelines — Reproducibility Package

收藏
Zenodo2026-07-23 更新2026-08-01 收录
官方服务:

资源简介:

v4 (2026-07-12): adds the LLM-judge-on-random-slice exploratory comparison (eval_llm_judge_on_m1.py + llm_judge_eval_on_m1.json): gpt-oss:20b and gemma4:e2b run under the B6 deterministic binary-grounding prompt (verbatim, temperature 0) over all 600 E11 items, scored against the 5-rater majority (recall 36.0%/34.7%, kappa 0.286/0.304, median ~2.4s/item). v3 content otherwise unchanged (de-identified R1-R5 rater files, sampling manifests, one-command derivation, pre-registrations, GPU-measured E4 incl. deployed bin-rule arm, injector-regime scripts/results). Restricted during peer review; CC BY 4.0 upon acceptance. Author metadata withheld for review.

提供机构:
Zenodo
创建时间:
2026-07-09
二维码
社区交流群
二维码
科研交流群
商业服务