fenyo/MonEspaceSante-FAQ-QA
收藏资源简介:
该数据集是一个法语问答对数据集,专门针对法国公共医疗服务Mon espace santé的常见问题(FAQ)构建。数据集包含614个问答对,分为两部分:89对来自官方FAQ的真实问答(real),以及525对通过qwen3.5模型生成的合成问答(synthetic),合成问答严格基于官方文本生成,不添加额外信息。数据集分为训练集(584对)和验证集(30对),每个问答对包含问题(question)、答案(answer)、来源(source,分为real或synthetic)和类别(category,对应FAQ的特定分类)。该数据集主要用于医疗健康领域的法语问答任务、文本生成和指令微调,特别适用于构建专注于Mon espace santé服务的封闭式问答助手。数据集规模较小(n<1K),适用于特定领域的微调,但不提供医疗建议,且知识截止于收集日期。
This dataset is a French question-answering (QA) pair dataset specifically constructed for frequently asked questions (FAQ) about the French public healthcare service *Mon espace santé*. It contains 614 QA pairs divided into two subsets: 89 real QA pairs sourced from the official FAQ, and 525 synthetic QA pairs generated via the Qwen3.5 model. The synthetic QA pairs are strictly generated based on the official text without adding any extraneous information. The dataset is split into a training set (584 pairs) and a validation set (30 pairs). Each QA pair includes four fields: question, answer, source (classified as either real or synthetic), and category (corresponding to the specific classification of the original FAQ). This dataset is primarily intended for French healthcare-domain QA tasks, text generation, and instruction fine-tuning, and is particularly suitable for building closed-domain QA assistants focused on the *Mon espace santé* service. The dataset has a small scale (n<1K), making it applicable for fine-tuning in specific domains, but it does not provide medical advice, and its knowledge cutoff is the date of collection.




