A Text-Level CONSORT 2025 Annotation Dataset of Randomized Controlled Trials
收藏资源简介:
This dataset supports sentence- and passage-level annotation mapping text from published randomized controlled trials (RCTs) to the items of the CONSORT 2025 reporting checklist. For this phase (this version), the dataset comprises only the source RCT PDFs and their metadata, the manuscripts that will be annotated. Annotation files (extracted passages + CONSORT item labels) are not included in this version and will be released in a subsequent version. RCTs were identified in PubMed (publication years 2020–2025), restricted to English-language primary reports of human trials with open-access full text. Manuscripts were selected by stratified random sampling across publication year, sample size (small n=20–199 vs. large n≥200), and journal impact (SCImago Q1–Q2 vs. Q3–Q4). In the planned annotation phase, two independent reviewers will extract exact text passages for each CONSORT 2025 item and label completeness as Complete, Partial, or Absent, with location and context notes. The dataset is intended as training/validation data for automated detection of CONSORT reporting elements. It accompanies the study protocol "Systematic Selection and Text-Level Data Extraction from Randomized Controlled Trials for CONSORT 2025 Checklist Item Mapping" (v1.22). Contents (this version): source open-access RCT PDFs and per-article metadata (title, journal, SCImago quartile, year, PMID, DOI, license, sample size). Annotation files will be added in a future version.



