MEDFACT
收藏资源简介:
MEDFACT是一个大规模的中文数据集,用于对大型语言模型(LLM)生成的医疗内容进行基于证据的医疗事实核查。该数据集包含1321个问题和7409个声明,反映了现实世界医疗场景的复杂性。数据集的构建包括收集医疗问题、使用LLM生成响应、将响应分解为原子声明、提取声明、检测声明的核查价值、检索证据和标注声明的真实性。数据集适用于医疗事实核查任务,旨在解决医疗信息在线传播中的不准确和误导问题。
MEDFACT is a large-scale Chinese dataset dedicated to evidence-based medical fact checking of medical content generated by large language models (LLMs). It contains 1,321 questions and 7,409 statements, which reflect the complexity of real-world medical scenarios. The construction of the dataset involves collecting medical questions, generating responses with LLMs, decomposing responses into atomic statements, extracting statements, detecting the verifiability of statements, retrieving evidence, and annotating the veracity of the statements. This dataset is applicable to medical fact checking tasks, aiming to address the issues of inaccuracy and misinformation in the online spread of medical information.




