Primary Literature Search and Deduplicated Analytical Corpus (2010–2025) for Cloud Computing in Healthcare: An OpEx Perspective
收藏资源简介:
This dataset serves as the primary analytical corpus for the systematic review titled "CLOUD COMPUTING IN HEALTHCARE: AN OPEX PERSPECTIVE" (F1000Research identifier: F1R-VER206246-A). Dataset Overview- Analytical Scope: 5,847 deduplicated bibliometric records (2010–2025) spanning OpenAlex and PubMed.- Pre-Screening Baseline: 6,764 total retrieved records (5,900 OpenAlex + 864 PubMed).- Deduplication Yield: 333 duplicates removed via R mergeDbSources() (DOI and fuzzy-title matching).- Quality Filters Applied: 1. Year restriction (2010–2025): -104 records 2. Language filter (English only): -109 records 3. Character threshold (Abstract > 50 characters): -371 records File Inventory1. `Final_Analytical_Corpus_5847.csv`: The complete, filtered dataset containing full citation metadata, OpenAlex Concept IDs (C2908647359 & C71924100), MeSH terms, abstract text, publication years, and source metrics.2. `P1_F1000Research_Anonymous.docx`: Anonymized supplementary manuscript file. Data Use & LicensingThis dataset is published under the Creative Commons Zero (CC0 1.0) Public Domain Dedication in full compliance with F1000Research Open Data Policy. le 1. PRISMA 2020 Literature Selection Flow Screening Stage Records (n) OpenAlex (concept IDs C2908647359 + C71924100, keyword-filtered) 5,900 PubMed (Title/Abstract + MeSH headings) 864 Total before deduplication 6,764 After mergeDbSources() — DOI and fuzzy-title deduplication 6,431 After publication-year filter (2010–2025) 6,327 After language filter (English only) 6,218 After abstract-availability filter (character length > 50) 5,847 Final analytical corpus 5,847



