Long-Form Medical Question Answering Benchmark
收藏资源简介:
Long-Form Medical Question Answering Benchmark是由Lavita AI和达特茅斯学院合作创建的一个公开可用基准数据集,专注于长格式医疗问答。该数据集包含1298条真实世界的消费者医疗问题,这些问题经过医学专家的注释和评估。数据集的创建过程包括用户查询的收集、去重、语义去重和质量检查。该数据集旨在评估大型语言模型在医疗领域的长格式回答生成能力,解决现有基准数据集在真实临床应用中的不足。
Long-Form Medical Question Answering Benchmark is a publicly available benchmark dataset co-created by Lavita AI and Dartmouth College, focusing on long-form medical question answering. It contains 1,298 real-world consumer medical questions that have been annotated and evaluated by medical experts. The dataset construction process includes the collection, deduplication, semantic deduplication, and quality inspection of user queries. This benchmark aims to evaluate the long-form answer generation capability of large language models in the medical domain, addressing the limitations of existing benchmark datasets in real-world clinical applications.




