Deduplicated bibliometric corpus: Large Language Models in healthcare simulation and non-technical skills training (2020-2026, 86,635 records)
收藏资源简介:
This dataset is the deduplicated bibliometric corpus underlying the manuscript "Large Language Models in Healthcare Simulation Education: A Bibliometric Analysis with AI-Assisted Screening" (submitted to PLOS Digital Health, 2026). It contains 86,635 unique scholarly records on the application of large language models (LLMs) to healthcare simulation and non-technical skills (NTS) training, retrieved from seven open-access databases (OpenAlex, PubMed, Europe PMC, Crossref, Semantic Scholar, CORE, DOAJ) covering January 2020 to March 2026. Schema (11 columns): id, doi, title, year, journal, source_db, citation_count, type, open_access, found_in_sources, abstract. Source-database breakdown of the 86,635 records: OpenAlex 46,826; Crossref 16,315; Europe PMC 14,162; Semantic Scholar 4,951; CORE 3,313; PubMed 1,062; DOAJ 6. From this corpus a sequential keyword funnel produced 830 candidate papers, screened by 83 independent AI agents (Claude Sonnet 4.6) to yield a verified corpus of 551 papers (Cohen's kappa pre-reconciliation = 0.86, post-reconciliation = 1.0). Complete screening decisions, per-paper AI rationales, and the full analysis pipeline are openly available in the linked GitHub repository.



