Reproducibility dataset for "Mapping the Research Landscape of Extrafloral Nectaries: A Comprehensive Bibliometric Analysis (1894–2026)
收藏资源简介:
This Zenodo deposit contains the complete reproducibility package for the bibliometric study "Mapping the Research Landscape of Extrafloral Nectaries: A Comprehensive Bibliometric Analysis (1894–2026)" by Basnet & Joshee, currently under review at Frontiers in Plant Science. Contents: Raw data exports from Web of Science (wos_part1.xls, wos_part2.xls; n = 1,700) and Scopus (scopus_export.csv; n = 1,291), retrieved February 2026 using the search string described in Methods Section 2.1 of the manuscript. Reproducibility script (run_pipeline.R): a single, version-controlled R script that takes the three raw exports as input and reproduces every step of the PRISMA screening flow — cross-database deduplication via bibliometrix::mergeDbSources(), manual screening using the four operational exclusion criteria (E1, E2, E3, E4) defined in Methods Section 2.2, and final dataset construction. Cleaned analytical dataset (EFN_combined_clean.RData): the final 1,279-document bibliometrixDB object on which all bibliometric analyses (annual production, Bradford's Law, Lotka's Law, h-index, country and institutional analyses, keyword co-occurrence, thematic mapping, collaboration network) reported in the manuscript were performed. Record-level screening log (screening_log.csv): the inclusion/exclusion decision for each of the 1,924 unique records that entered manual screening, with the assigned exclusion code (E1, E3, or E4) for each of the 645 excluded records. Duplicate log (duplicates_log.csv): the 1,067 cross-database duplicates removed at the deduplication step, with the kept-vs-removed mapping. PRISMA summary (prisma_summary.txt): a plain-text summary of the four PRISMA counts (2,991 → 1,067 → 1,924 → 645 → 1,279). PRISMA flow: 2,991 records identified (Web of Science n = 1,700; Scopus n = 1,291) → 1,067 cross-database duplicates removed → 1,924 unique records screened → 645 records manually excluded (E1: floral-nectary-only studies; E3: tangential ant ecology; E4: residual within-database duplicates) → 1,279 documents in the final analytical dataset. Software requirements: R ≥ 4.0 with bibliometrix ≥ 4.0, readxl, dplyr, stringr, and stringdist; Python ≥ 3.10 with matplotlib, pandas, numpy, and scipy. All package versions are documented in the README. How to use: Place the three raw files in a raw_data/ directory, then run Rscript run_pipeline.R from the project root. The script will create the outputs/ directory and write the cleaned dataset, screening log, duplicates log, and PRISMA summary. Detailed instructions are in the README. Citation: If you use this dataset, please cite the associated manuscript (Basnet & Joshee, 2026, Frontiers in Plant Science) and this Zenodo deposit.



