遇见数据集

Synthetic Clinical Communication for Healthcare NLP: Datasets and Generation Code for Thirteen Application Case Studies

收藏
Zenodo2026-08-06 更新2026-08-13 收录
官方服务:

资源简介:

Reproducibility deposit accompanying the survey Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies (Apartsin and Aperstein). The paper studies how models trained on LLM-generated synthetic clinical communication perform on downstream clinical NLP tasks, grounded in thirteen application case studies spanning patient-facing narratives, clinician documentation, patient-portal messaging, and pre-hospital and handover communication. Contents. Case-study datasets and documentation (synthpatient_casestudies.zip): one folder per application study, each containing the LLM-generated synthetic data (or, for datasets above the size cap, a 200-line sample recording the full size and row count), the LLM data-generation script or notebook, and a README.md describing the clinical motivation, generation protocol, data files, models trained and compared, and results. An index README.md maps each folder to its section in the paper. Paper snapshot (synthpatient_paper.zip): the survey as paper.html with its figures, plus single-column and two-column (arXiv-style) renderings in PDF and Microsoft Word. Not included. Third-party and credentialed corpora used as seeds (for example MIMIC-IV-Ext-CDS on PhysioNet and the Symptom-Disease Prediction Dataset on Mendeley) are not redistributed; each study README gives retrieval instructions. Model weights, presentations, and audio are not redistributed. All shared data is synthetic (LLM-generated). Redistribute each dataset only as its source repository permits (see each folder's Notes / license section).

提供机构:
Zenodo
创建时间:
2026-08-06
二维码
社区交流群
二维码
科研交流群
商业服务