PraxySante/Qwen3-ASR-PostTrain-Complete-Medical-French-Full
收藏资源简介:
该数据集是一个多源的法语医学文本集合,包含来自多个不同来源的文本数据,例如药物警戒数据库(bdpm)、医疗转录(nemo_transcriptions)、患者病历(parcomed、parcomed_research)、医学文献(hitz_medical)、法语维基百科(wikipedia_fr)、网络爬取语料(c4_fr、fineweb2_fr)、对话(claire_dialogue)、字幕(open_subtitles_fr)等。每个样本包含文本内容和来源标签,总样本数超过千万,主要用于训练和评估法语自然语言处理模型,尤其是医学领域。
This dataset is a multi-source collection of French medical texts, comprising data from various sources such as pharmacovigilance database (bdpm), medical transcriptions (nemo_transcriptions), patient records (parcomed, parcomed_research), medical literature (hitz_medical), French Wikipedia (wikipedia_fr), web-crawled corpora (c4_fr, fineweb2_fr), dialogues (claire_dialogue), subtitles (open_subtitles_fr), and more. Each sample consists of a text field and a source label, with over 10 million examples in total, intended for training and evaluating French NLP models, particularly in the medical domain.



