遇见数据集

CiteMe Citation Verification Benchmark v1

收藏
Zenodo2026-07-26 更新2026-08-13 收录
官方服务:

资源简介:

A frozen multilingual benchmark for evaluating citation- and reference-verification systems. It contains 340 labeled references across 116 evaluation cases, with ground truth known by construction rather than inferred from user submissions. Ground-truth class Count Construction EXISTS_CORRECT 160 Real works confirmed by DOI and title in OpenAlex and Crossref EXISTS_CORRUPTED 100 Real works with exactly one controlled metadata corruption FABRICATED 80 Constructed references whose invented titles had no close OpenAlex or Crossref match at build time Coverage. English (170 references), Portuguese (102), Spanish (51), German (17); APA (146), Vancouver (90), Harvard (66), MLA (24), ABNT (14); 102 single-reference cases and 14 bibliography bundles, 23 of which use deliberately untidy hand-typed formatting. Construction and provenance. Correct references start from public OpenAlex metadata and are confirmed against Crossref before inclusion. Corrupted references receive exactly one deterministic mutation to the title, year, author, journal, DOI, or pages/volume field. Fabricated references are constructed from synthetic titles and public bibliographic strings, then rejected if OpenAlex or Crossref returned a close title match at build time. The build seed is 20260723. Because OpenAlex and Crossref evolve, the frozen files and their SHA-256 checksums, not a future API rerun, are the reproducibility anchor for v1. Synthetic-data notice. Fabricated rows may combine public contributor and venue strings with invented titles. They are synthetic negative test cases and must never be interpreted as claims that a named person authored, or a venue published, the constructed work. Limitations. This is an evaluation corpus, not a prevalence sample of academic writing. Non-existence checks describe the OpenAlex/Crossref snapshot at build time. The 29 separately curated legitimate but unindexed works used in CiteMe's 369-reference product evaluation are not part of corpus-v1. Future corrections or additions will be released as a new version; v1 files will not be edited in place. The corpus contains no user submissions, bibliographies, documents, product analytics, or other user data. Canonical page: https://citeme.app/datasets/citation-verification-benchmark

提供机构:
Zenodo
创建时间:
2026-07-26
二维码
社区交流群
二维码
科研交流群
商业服务