遇见数据集

High sensitivity title and abstract screening pipeline

收藏
Zenodo2025-07-20 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains the data of a study that developed and evaluated an automated pipeline for screening titles and abstracts in systematic reviews. The primary goal of this research was to create a highly sensitive and efficient screening process using fine-tuned, open-source large language models (LLMs) combined with a PICOS (Population, Intervention/Exposure, Comparison, Outcome, Study Design) framework. Dataset Contents The dataset is organized into four main folders, each corresponding to one of the four systematic reviews used in the study: Study 1: A systematic review and meta-analysis of randomized controlled trials. Study 2: A systematic review and meta-analysis of cohort studies. Study 3: A systematic review and meta-analysis of cohort studies or studies with data derived from randomized controlled trials. Study 4: A systematic review of observational studies. Each study folder contains a set of CSV files that document the screening process at different stages: Total articles: This file lists all the articles retrieved from the initial literature search for the respective systematic review. Study design screen: This file contains the articles that were excluded by the automated pipeline based on their study design. PICO screen: This file includes the articles that were excluded after applying the simplified PICOS criteria. Final included: This file lists the articles that were ultimately included for full-text screening after the automated pipeline was applied. Methods The data was generated using a multi-step automated pipeline: PICOS Summarization: A fine-tuned Mistral-Nemo-Instruct-2407 model was used to generate structured PICOS summaries from the titles and abstracts of the retrieved articles. Exclusion of Irrelevant Study Designs: A fine-tuned Qwen3-14B model was employed to exclude non-human studies, reviews, and other non-primary literature. Stepwise PICOS-based Screening: The same Qwen3-14B model, trained on a stepwise screening strategy, was used to apply the simplified PICOS inclusion criteria to identify relevant articles. The performance of this pipeline was evaluated by comparing its results to the gold standard of included articles from the original published systematic reviews. The dataset provides the complete outputs of this automated process, allowing for full transparency and reproducibility of the study's findings. This data can be used to validate the results, explore the performance of the models in detail, and serve as a foundation for future research in AI-assisted evidence synthesis.

提供机构:
Zenodo
创建时间:
2025-07-20
二维码
社区交流群
二维码
科研交流群
商业服务