遇见数据集

Closing the Domain Gap: An Author-Labeled Computer Science Dataset for Title and Abstract Screening Automation

收藏
Zenodo2026-07-25 更新2026-08-02 收录
官方服务:

资源简介:

A pooled, criteria-conditioned corpus of 22 systematic literature reviews (SLRs) spanning computer science, artificial intelligence, and software engineering, built to support research on automating the title/abstract screening stage of systematic reviews. The dataset contains 16,763 candidate paper records (2,732 included, 14,031 excluded; 16.3% overall inclusion rate), each labeled with a binary include/exclude decision by the original review's authors. Unlike single-review screening datasets, each paper record here is paired with its originating review's own inclusion and exclusion criteria, title, keywords, and research question — allowing screening to be modeled as conditioned on a specific review's rules rather than as a fixed topic classification task. Per-review inclusion rates vary substantially (approximately 3.4%–64.5%), making the dataset well suited to class-imbalance research at both the dataset-wide and per-review level. Two files are provided: original_slr_dataset.xlsx, containing full record-level data (title, abstract, authors, venue, DOI, database source, and criteria fields) linked to its source review via a source_file identifier; and overall_per_slr_summary.csv, a review-level summary of paper counts and inclusion rates across all 22 reviews. Intended applications include automated screening model training and benchmarking (classical ML, deep learning, and transformer-based approaches), embedding/representation comparison studies, class-imbalance handling research, cross-review generalization and transfer learning experiments, and criteria-conditioned NLP/entailment-style modeling.

提供机构:
Zenodo
创建时间:
2026-07-25
二维码
社区交流群
二维码
科研交流群
商业服务