Dataset for: SESR-Eval: Dataset to Evaluate LLMs in the Screening Process of Systematic Reviews Creators
收藏资源简介:
Introduction This is a dataset for: "SESR-Eval: Dataset to Evaluate LLMs in the Screening Process of Systematic Reviews". Folder structure data The `data`-folder contains: - Initial replication package selection (`1-replication-package-selection`) - Inter-rater reliablity agreement for replication package selection (`2-replication-package-selection-reliability-agreement`) - Processed replication packages (`3-processed-data`) - Replication packages are omitted due to size constraints, but are downloadable via provided links - LLM results (`4-llm-results`) - The SESR-Eval dataset (`sesr-eval-dataset`) See: `data/sesr-eval-dataset/README.md` documentation The `documentation`-folder contains miscellaneous documentation for the study. experiments The `experiments`-folder contains the LLM experiment source code. How to run the benchmarks? 1. Install Python 3 2. Run `python3 -m venv venv` 3. Run `source venv/bin/acticate` 4. Run `pip install -r requirements.txt` 5. Copy `.env.example` to `.env` 6. Obtain: 1. Dataset (see data/sesr-eval-dataset/README.md) 2. OpenAI API key 3. Openrouter API key (if you wish to run other models than OpenAI) 7. Run: `./run_experiments.sh` Requirements - Python 3 Scopus API usage The data was downloaded from Scopus API between January 1 and 25 April, 2025 via http://api.elsevier.com and http://www.scopus.com. License The replication package is licensed with the CC-BY-ND 4.0 license. Each dataset secondary study has their own license. However, Elsevier has their own terms and conditions regarding the use of our research data: ---- This work uses data that was downloaded from Scopus API between Jan 1 and Apr 24, 2025 via http://api.elsevier.com and http://www.scopus.com. Elsevier allows access to the Scopus APIs in support of academic research for researchers affiliated with a Scopus subscribing institution. The end product here is a scholarly published work, that utilizes publications in Scopus for our research effort. We want to publish a scholarly work regarding Scopus data relationships. The data downloaded from Scopus API, for our work, is published to make work follow the practices of open science. It also makes possible to reproduce our work's results. Elsevier allows this use case under the following conditions, which our work meets: - The research is for non-commercial, academic purposes only. - The research is performed by approved representative of the applying institution. - The research is limited to the scope of Software engineering (SE) - we are not mining the entire Scopus dataset. - The retention of original research dataset is limited to archival purposes and reproduction of the research results. - Public sharing of data for purpose of reproducibility with a specific party is permissible upon written request and explicit written approval. - Scopus has been identified as the data source as described in the Scopus Attribution Guide. - If the user is a bibliometrician doing work outside this use case, they contact Elsevier's International Center for the Study of Research. The data is not displayed in a website or in a public forum outisde of the output format of the scholarly published work. The data is only stored in Zenodo, in this replication package.
## 数据集介绍 本数据集用于:“SESR-Eval:用于评估大语言模型(Large Language Model)在系统评价筛查流程中表现的数据集”。 ## 文件夹结构 ### data 文件夹 `data` 文件夹包含以下内容: - 初始复现包遴选(`1-replication-package-selection`) - 复现包遴选的评定者间一致性信度(`2-replication-package-selection-reliability-agreement`) - 处理后的复现包(`3-processed-data`) - 复现包因体积限制未纳入本包,但可通过提供的链接下载 - 大语言模型结果(`4-llm-results`) - SESR-Eval 数据集(`sesr-eval-dataset`) 详见:`data/sesr-eval-dataset/README.md` ### documentation 文件夹 `documentation` 文件夹包含本研究的各类辅助文档。 ### experiments 文件夹 `experiments` 文件夹包含大语言模型实验的源代码。 ## 如何运行基准测试? 1. 安装 Python 3 2. 执行 `python3 -m venv venv` 创建虚拟环境 3. 执行 `source venv/bin/activate` 激活虚拟环境(注:原文本存在拼写笔误,已修正为正确命令) 4. 执行 `pip install -r requirements.txt` 安装依赖包 5. 将 `.env.example` 复制为 `.env` 6. 获取以下资源: 1. 数据集(详见 `data/sesr-eval-dataset/README.md`) 2. OpenAI API 密钥 3. Openrouter API 密钥(如需运行 OpenAI 以外的模型) 7. 执行 `./run_experiments.sh` ## 依赖要求 - Python 3 ## Scopus API 使用说明 本数据集于2025年1月1日至4月25日通过 http://api.elsevier.com 和 http://www.scopus.com 从 Scopus API 下载获取。 ## 授权协议 本复现包采用 CC-BY-ND 4.0 协议授权。各数据集二级研究拥有各自的授权协议。不过爱思唯尔(Elsevier)针对本研究数据的使用另有条款与细则: 本研究使用的数据于2025年1月1日至4月24日通过 http://api.elsevier.com 和 http://www.scopus.com 从 Scopus API 下载获取。 爱思唯尔允许 Scopus 订阅机构的研究人员使用 Scopus API 开展学术研究。 本研究的最终成果为学术出版物,其研究过程中使用了 Scopus 收录的文献。我们计划发表一篇关于 Scopus 数据关联的学术论文。 本研究从 Scopus API 下载的数据已公开,以遵循开放科学实践,并确保本研究成果可被复现。 爱思唯尔允许本研究使用 Scopus 数据,且本研究符合以下使用条件: - 本研究仅用于非商业性学术用途 - 研究由申请机构的获批代表开展 - 研究范围限定于软件工程(Software Engineering,SE)领域——我们未对整个 Scopus 数据集进行挖掘 - 原始研究数据集仅用于存档与研究成果复现 - 为实现成果复现而向特定方公开共享数据,需经书面申请并获得明确书面批准后方可进行 - 已按照《Scopus 归属指南》明确标注 Scopus 为数据源 - 若使用者为文献计量研究者且开展的工作超出本使用场景,请联系爱思唯尔国际研究中心(Elsevier's International Center for the Study of Research) 本数据仅以学术出版物的输出格式呈现,未在网站或公共论坛公开发布。本复现包中的数据仅存储于 Zenodo 平台。



