doctolib-lab/finemed-fr
收藏资源简介:
FineMed-fr 是一个大规模、公开可用的法语医学文本语料库,用于语言模型预训练。它包含来自三个异构开放网络源(FineWeb-2、FinePDFs 和 FineWiki)的医学内容,总计 21.1M 文档和 19.2B 单词,涵盖了真实世界的医学写作。每个文档都沿三个维度进行标注:子领域(分为 15 个医学子领域,如生物医学与临床写作、面向消费者的材料)、教育质量(基于 FineWeb-Edu 改编的 0-5 分评分)和医学术语密度(通过提取的医学术语字符覆盖比例衡量)。数据集以未过滤形式发布,用户可以根据注释列自定义阈值以适应任务需求。
FineMed-fr is a large, openly available corpus of French medical text for language-model pretraining: 21.1M documents and 19.2B words of real-world medical writing, annotated along several quality axes. The corpus is drawn from three heterogeneous open-web sources (FineWeb-2, FinePDFs, and FineWiki), which together provide the scale, source diversity, and stylistic range that curated medical corpora often lack. We keep only the French medical content, then label every surviving document along three axes: Subdomain (which of 15 medical subdomains the document belongs to), Educational quality (how instructive the document is for medical education, scored 0–5), and Medical-term density (the richness of medical terminology, measured as the fraction of characters that fall inside extracted medical-term spans). The corpus is released unfiltered so users can set their own thresholds on the annotation columns to fit their task.




