遇见数据集

Public Dataset for Offline Medical Speech Recognition Evaluation: The French Sovereign Clinical Corpus

收藏
Zenodo2026-03-16 更新2026-05-26 收录
官方服务:

资源简介:

The rapid integration of Artificial Intelligence in healthcare has predominantly relied on cloud-based architectures, raising severe concerns regarding patient confidentiality and data sovereignty. To address the critical lack of transparent training resources for offline AI, we present the French Sovereign Clinical Corpus. This comprehensive, open-source dataset is explicitly designed to train, evaluate, and benchmark local, air-gapped Speech-to-Text models and Large Language Models (LLMs) in the French medical domain. The dataset comprises over 1,000 highly curated Question/Answer instruction pairs alongside diverse simulated clinical audio baseline data. By utilizing a hybrid methodology of anonymized structural templates and procedural synthetic generation, the corpus entirely circumvents Protected Health Information (PHI) risks. It provides a standardized "Ground Truth" for researchers and IT directors aiming to deploy Sovereign AI solutions like LucioleScribe, ensuring compliance with strict legal frameworks such as GDPR and HIPAA without sacrificing state-of-the-art transcription accuracy.

提供机构:
Zenodo
创建时间:
2026-03-16
二维码
社区交流群
二维码
科研交流群
商业服务