xPeerd Epistemic Benchmark: Re3-Sci2 (v4 parallel results)
收藏资源简介:
This dataset provides a curated collection of 1,108 validated evaluation cases for assessing the performance of AI-driven peer-review systems. It maps high-quality human ground truth from the Re³-Sci 2.0 archive (F1000Research v1 manuscripts) against synthetic reviews generated by the xPeer simulation engine. Methodology & Leakage Prevention To ensure a rigorous evaluation, this benchmark was constructed following a strict leakage-safe protocol: Isolation: Manuscript text and metadata were extracted from the source documents and stored independently from human reviews. Blind Submission: The xPeer simulation engine was provided only with the manuscript content and metadata. No human recommendations, decisions, or review texts were included in the API payloads. Post-Hoc Joining: The human 'Ground Truth' was merged with the synthetic 'xPeer Simulation' only after the simulation outputs were frozen and persisted to disk. Data Content The primary archive xpeerd_benchmark_study_2026_v1.0.0.zip contains: benchmark_id: A stable, hashed identifier for each case. manuscript: Title, abstract, and full-text content of the F1000Research v1 document. human_reviews: Original reports from at least two human reviewers, including specific recommendations (e.g., approve, approve-with-reservations). xpeer_simulation: Four structured synthetic outputs: Reviewer1 Report Reviewer2 Report Editorial_summary Recommendation (The final simulated decision) Sanitization All records have been programmatically scrubbed to remove sensitive transport metadata, including API keys, internal URLs, and endpoint specific headers, ensuring compliance with security best practices for public data sharing. Usage This dataset is intended for researchers in Scientometrics, NLP, and AI Ethics to study the alignment between LLM-based editorial agents and human expert judgment.



