A French OSCE Dialogue Dataset for Clinical Training
收藏资源简介:
The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which consist of brief scenario-driven simulations of doctor-patient interactions. However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs). To address this gap, we introduce a French OSCE dialogue dataset comprising 294 student-patient training interactions. Component Count OSCE stations recorded 29 Recorded dialogues 294 Automatically transcribed dialogues 294 Corrected transcriptions 26 Total audio duration (hours) 30 Dataset structure Size: 8 MB The dataset is organized into two main directories: cases/: contains the description and materials associated with each OSCE station. recorded/: contains the recorded interactions. Each subdirectory corresponds to a station identifier matching the one used in cases/. Each station folder contains: the automatically generated speaker-diarized transcription, the manually corrected transcription when available. This resource is intended to support research in medical communication, spoken dialogue systems, clinical speech processing, and the development of virtual patients for medical education. Dataset access This dataset is currently distributed under restricted access due to the inclusion of human participant data. Access requests can be submitted directly through the Zenodo record and will be evaluated on a case-by-case basis for research and educational purposes. Additional information For a detailed description of the data collection protocol, annotation process, recording setup, and automatic transcription pipeline, please refer to the associated publication "A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training". For any questions regarding the dataset or access conditions, please contact doria.bonzi@loria.fr.



