遇见数据集

PraxySante/t5-semantic-finetuning-dataset-fr-medical

收藏
Hugging Face2026-05-24 更新2026-05-31 收录
官方服务:

资源简介:

T5语义微调数据集 - 法语医学ASR校正是一个用于法语医学自动语音识别(ASR)语义校正的监督错误校正对数据集。该数据集包含248,609个训练对以及验证和评估分割,专门设计用于微调T5模型在医学ASR转录校正任务上的性能。数据格式为CSV,包含input、target、source三列,其中输入格式为corriger l erreur asr: <错误转录>,目标格式为<校正后的转录>。错误类型涵盖语义错误(如药物名称混淆、医学术语替换)、词汇错误(如音似词替换)和语法错误(如医学上下文中的一致性、代词、限定词错误)。数据集来源多样,包括Wadiaa nemo语义增强、药物混淆对、相似词汇药物替换、已验证的错误校正对、一般语义校正、医学分类术语混淆、完整验证对、医学实体混淆和临床案例标记校正等。该数据集适用于自然语言处理任务,特别是文本到文本生成,在医学领域具有实际应用价值。

The T5 Semantic Fine-Tuning Dataset - French Medical ASR Correction is a supervised error correction pairs dataset for semantic correction of French medical automatic speech recognition (ASR) transcriptions. It contains 248,609 training pairs along with validation and evaluation splits, specifically designed for fine-tuning T5 models on medical ASR transcription correction tasks. The data is in CSV format with columns: input, target, source, where the input format is corriger l erreur asr: <erroneous transcription> and the target format is <corrected transcription>. Error types include semantic errors (e.g., drug name confusions, medical term substitutions), vocabulary errors (e.g., phonetically similar word replacements), and grammatical errors (e.g., agreement, pronoun, determiner errors in medical context). The dataset is sourced from various components such as Wadiaa nemo semantic augmented, drugs confusables, drugs similar words, aggregated minimal, semantic, CIM-11, aggregated pairs, QUAERO entities, and CAS tokens. It is suitable for natural language processing tasks, particularly text-to-text generation, with practical applications in the medical domain.

提供机构:
PraxySante
二维码
社区交流群
二维码
科研交流群
商业服务