遇见数据集

FreshState: An Endpoint-Validated, Age-Matched Benchmark for Stale Evidence in Web-Augmented Language Models

收藏
Zenodo2026-05-28 更新2026-05-26 收录
官方服务:

资源简介:

FreshState v2 is an endpoint-validated, age-matched, source-unique benchmark and reusable resource for stale-evidence prediction in retrieval-augmented language models. The released Task 1 evaluation set contains 524 examples (262 stale + 262 fresh): apartment 59/59, software 63/63, and PyPI 140/140. The construction preserves exact within-domain age matching (KS D = 0, p = 1.0; mean age 9.77 days in both classes), enforces within-class source uniqueness, and has zero within-domain source overlap between the stale and fresh classes. Candidate-change seeds are collected through two evidence paths. For the web domains, Craigslist and GitHub candidate events are emitted by prospective daily monitoring. For PyPI, candidate events are reconstructed from structured registry history over a checkpoint window. Released examples in all domains are retained only after domain-appropriate endpoint validation. Software final answer-bearing values are reconstructed from publicly visible GitHub Releases API history under an API-canonical stable-tag definition: canonical_stable_tag(releases, D) is the latest non-draft, non-prerelease release with published_at <= D 23:59:59 UTC. PyPI answer-bearing values are defined as the latest PEP 440-parseable non-prerelease, non-development release recorded in PyPI registry metadata at the checkpoint date, retaining releases marked as yanked. The bundled PyPI provenance snapshot supports offline verification of all retained PyPI examples; a yanked-excluding alternative would affect 4 of the 280 retained PyPI Task 1 examples. v2 applies three corrections identified during final eligibility audits of v1 (DOI: 10.5281/zenodo.20337401): endpoint-supported fresh candidates at cached and query times; within-domain source uniqueness on both stale and fresh sides; and API-canonical software endpoint validation replacing noisy HTML-extractor value assignment. On the released Task 1 set, the query-at-aligned GPT-4o-mini verifier remains near chance (balanced accuracy 50.8, macro-F1 50.7, AUROC 50.8). The artifact also includes a final Option D-aligned controlled snippet-swap diagnostic for GPT-4o and GPT-4o-mini, together with saved outputs and offline reproduction scripts. The artifact bundles the evaluation set, naive ablation set, candidate-change seeds, endpoint-support and selection reports, minimal GitHub Releases API provenance, apartment first-observed evidence, minimal PyPI registry-history provenance, validation logs, prompt templates, saved baseline and verifier outputs, final snippet-swap outputs, and scripts for offline reconstruction and table reproduction. Project-authored source code and reproduction scripts are released under the MIT License. FreshState-authored annotations, labels, endpoint-support verdicts, selection reports, audit judgments, prompt templates, result files, evaluation-set organization, and documentation are released under Creative Commons Attribution 4.0 International. Source-derived factual fields are included solely as provenance required to reproduce the benchmark; FreshState does not assert ownership over or relicense underlying third-party platform content represented in those fields. Canonical repository: https://github.com/kelu-look/freshstate

提供机构:
Zenodo
创建时间:
2026-05-22
二维码
社区交流群
二维码
科研交流群
商业服务