xPeerd Epistemic Benchmark: Re\u00b3-Sci 2.0 (v3.1 parallel results)
收藏资源简介:
xPeerd Benchmark Dataset Scholarly peer review is overwhelmed by surging submissions and reviewer shortages; existing pressures now intensified by generative AI. While AI promises to streamline editorial workflows, lagging governance risks compromising academic integrity. Although AI assistants accelerate synthesis and expand coverage, they struggle with complex reasoning, hallucinate references, and foster user over-reliance. Concurrently, large language models are rapidly advancing as zero-shot reasoners, expanding into complex, multimodal applications. To stabilize the publishing ecosystem, the scientific community must adopt a pragmatic approach. Such approach should Leverage AI for efficiency while mandating strict human oversight, transparency, and accountability. This dataset provides a systematic evaluation of xPeerd peer-review simulations against human gold-standard peer-review reports. It is derived from the Re3-Sci 2.0 corpus and contains 1000 deterministic benchmark cases from F1000Research (v1 manuscripts). Epistemic Scoring Framework Each human and simulated peer-review is evaluated using a deterministic, multi-dimensional scoring heuristic (0-100) comprising: Human Issue Coverage (35%): Similarity-based recall of concerns raised by human reviewers. Manuscript Grounding (15%): Overlap with specific manuscript tokens and structural markers (Figures/Tables). Actionability (15%): Density and detail of concrete revision recommendations. Epistemic Breadth (10%): Coverage across key review dimensions (Methods, Stats, Ethics, etc.). Calibration Alignment (10%): Correspondence of hedging and certainty markers compared to human baseline. Reasoning Density (10%): Frequency of logical connectors and evidence-based justifications. Non-redundancy (5%): Internal diversity of the generated review text.



