Synthetic Clinical Communication for Healthcare NLP: Datasets and Generation Code for Ten Application Case Studies
收藏资源简介:
Reproducibility deposit accompanying the survey Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies (Apartsin and Aperstein). The paper studies how models trained on LLM-generated synthetic clinical communication perform on downstream clinical NLP tasks, grounded in ten application case studies spanning patient-facing narratives, clinician documentation, patient-portal messaging, and pre-hospital and handover communication. This is a data-only deposit. It contains synthpatient_casestudies.zip: one folder per application study, each holding the LLM-generated synthetic data (or, for datasets above the size cap, a 200-line sample recording the full size and row count), the LLM data-generation script or notebook, and a README.md describing the clinical motivation, generation protocol, data files, models trained and compared, and results. An index README.md maps each folder to its section in the paper. Not included. The survey paper itself is not deposited here; it is published as a web document (see the related identifiers). Third-party and credentialed corpora used as seeds (for example MIMIC-IV-Ext-CDS on PhysioNet and the Symptom-Disease Prediction Dataset on Mendeley) are not redistributed; each study README gives retrieval instructions. Model weights, presentations, and audio are not redistributed. All shared data is synthetic (LLM-generated). Redistribute each dataset only as its source repository permits (see each folder's Notes / license section).



