遇见数据集

Reproducibility dataset for "Mapping the Research Landscape of Extrafloral Nectaries: A Comprehensive Bibliometric Analysis (1894–2026)

收藏
Zenodo2026-05-07 更新2026-05-26 收录
官方服务:

资源简介:

This Zenodo deposit contains the complete reproducibility package for the bibliometric study "Mapping the Research Landscape of Extrafloral Nectaries: A Comprehensive Bibliometric Analysis (1894–2026)" by Basnet & Joshee, currently under review at Frontiers in Plant Science. Contents: Raw data exports from Web of Science (wos_part1.xls, wos_part2.xls; n = 1,700) and Scopus (scopus_export.csv; n = 1,291), retrieved February 2026 using the search string described in Methods Section 2.1 of the manuscript. Reproducibility script (run_pipeline.R): a single, version-controlled R script that takes the three raw exports as input and reproduces every step of the PRISMA screening flow — cross-database deduplication via bibliometrix::mergeDbSources(), manual screening using the four operational exclusion criteria (E1, E2, E3, E4) defined in Methods Section 2.2, and final dataset construction. Cleaned analytical dataset (EFN_combined_clean.RData): the final 1,279-document bibliometrixDB object on which all bibliometric analyses (annual production, Bradford's Law, Lotka's Law, h-index, country and institutional analyses, keyword co-occurrence, thematic mapping, collaboration network) reported in the manuscript were performed. Record-level screening log (screening_log.csv): the inclusion/exclusion decision for each of the 1,924 unique records that entered manual screening, with the assigned exclusion code (E1, E3, or E4) for each of the 645 excluded records. Duplicate log (duplicates_log.csv): the 1,067 cross-database duplicates removed at the deduplication step, with the kept-vs-removed mapping. PRISMA summary (prisma_summary.txt): a plain-text summary of the four PRISMA counts (2,991 → 1,067 → 1,924 → 645 → 1,279). PRISMA flow: 2,991 records identified (Web of Science n = 1,700; Scopus n = 1,291) → 1,067 cross-database duplicates removed → 1,924 unique records screened → 645 records manually excluded (E1: floral-nectary-only studies; E3: tangential ant ecology; E4: residual within-database duplicates) → 1,279 documents in the final analytical dataset. Software requirements: R ≥ 4.0 with bibliometrix ≥ 4.0, readxl, dplyr, stringr, and stringdist; Python ≥ 3.10 with matplotlib, pandas, numpy, and scipy. All package versions are documented in the README. How to use: Place the three raw files in a raw_data/ directory, then run Rscript run_pipeline.R from the project root. The script will create the outputs/ directory and write the cleaned dataset, screening log, duplicates log, and PRISMA summary. Detailed instructions are in the README. Citation: If you use this dataset, please cite the associated manuscript (Basnet & Joshee, 2026, Frontiers in Plant Science) and this Zenodo deposit.

本Zenodo存档包含Basnet与Joshee所著文献计量学研究《花外蜜腺研究图景全景:1894–2026年综合文献计量分析》的完整可复现套件,该论文目前已提交至《Frontiers in Plant Science》(植物科学前沿)期刊待审。 ### 内容清单 1. 原始数据导出文件:来自Web of Science的wos_part1.xls、wos_part2.xls(共1700条记录),以及Scopus的scopus_export.csv(共1291条记录)。所有数据均于2026年2月根据手稿方法章节2.1所述检索式获取。 2. 可复现性脚本(run_pipeline.R):单份受版本控制的R脚本,以三份原始数据导出文件为输入,可复现PRISMA筛选流程的全部步骤——包括通过bibliometrix::mergeDbSources()实现跨数据库去重、采用手稿方法章节2.2定义的四项操作排除标准(E1、E2、E3、E4)进行人工筛选,以及最终数据集构建。 3. 清洗后分析数据集(EFN_combined_clean.RData):最终包含1279篇文献的bibliometrixDB对象,为手稿中所有文献计量分析(年度产出、布拉德福定律、洛特卡定律、h指数、国家与机构分析、关键词共现、主题映射、合作网络分析)的分析基础。 4. 记录级筛选日志(screening_log.csv):纳入/排除所有进入人工筛选的1924条唯一记录的决策记录,其中645条被排除的记录均标注了对应排除代码(E1、E3或E4)。 5. 重复记录日志(duplicates_log.csv):去重步骤中移除的1067条跨数据库重复记录,包含保留记录与移除记录的对应映射关系。 6. PRISMA汇总表(prisma_summary.txt):以纯文本形式呈现四项PRISMA统计计数:2991 → 1067 → 1924 → 645 → 1279。 ### PRISMA流程图说明 共识别到2991条记录(Web of Science:1700条;Scopus:1291条)→ 移除1067条跨数据库重复记录 → 对1924条唯一记录开展人工筛选 → 645条记录经人工排除(E1:仅研究花蜜腺的文献;E3:相关蚂蚁生态学边缘研究;E4:数据库内部残留重复记录)→ 最终分析数据集包含1279篇文献。 ### 软件依赖要求 R语言版本≥4.0,需安装bibliometrix≥4.0、readxl、dplyr、stringr及stringdist包;Python语言版本≥3.10,需安装matplotlib、pandas、numpy及scipy包。所有包的版本信息均已在README文件中记录。 ### 使用方法 将三份原始数据文件放置于raw_data/目录下,随后从项目根目录执行命令`Rscript run_pipeline.R`。脚本将自动创建outputs/目录,并写入清洗后数据集、筛选日志、重复记录日志及PRISMA汇总文件。详细操作说明见README文件。 ### 引用说明 若使用本数据集,请引用相关研究手稿(Basnet & Joshee, 2026, Frontiers in Plant Science)及本Zenodo存档包。

提供机构:
Zenodo
创建时间:
2026-05-07
二维码
社区交流群
二维码
科研交流群
商业服务