PraxySante/t5-posttraining-semantic-dataset-fr-medical
收藏资源简介:
Qwen后训练数据集(Enrichi)是一个法语文本生成数据集,专门用于医疗领域,支持Qwen模型的后训练。数据集规模在10万到100万条数据之间,包含来自多个来源的文本,如PARTAGES/PARHAF、QUAERO、CAS、MORFITT、PARCOMED、DEFT2021、HiTZ Medical FR、Wikipedia FR和C4 FR。数据以JSONL格式存储,主要字段为text,旨在丰富模型在法语医疗文本生成任务中的能力。
The Qwen Post-training Dataset (Enrichi) is a French text generation dataset tailored for the medical domain, designed to facilitate post-training of Qwen models. The dataset comprises between 100,000 and 1,000,000 data entries, incorporating textual data from multiple sources including PARTAGES/PARHAF, QUAERO, CAS, MORFITT, PARCOMED, DEFT2021, HiTZ Medical FR, Wikipedia FR, and C4 FR. The data is stored in JSONL format, with the core field being "text". This dataset aims to enhance the model's capabilities in French medical text generation tasks.



