Med-MMHL
收藏资源简介:
Med-MMHL是由弗吉尼亚理工大学创建的一个多模态医疗领域错误信息检测数据集,旨在解决现有数据集忽视视觉信息、仅关注COVID-19相关错误信息以及忽略大型语言模型生成错误信息的问题。该数据集不仅包含人为生成的错误信息,还涵盖了如ChatGPT等大型语言模型生成的错误信息,涉及15种疾病,数据来源于新闻和推文。创建过程中,通过爬取文本和相关图像,确保了数据的多模态性。该数据集的应用领域广泛,旨在提升医疗领域错误信息的检测能力,特别是在区分人为和语言模型生成的错误信息方面。
Med-MMHL is a multimodal medical misinformation detection dataset developed by Virginia Tech. It aims to address the core limitations of existing datasets, including neglect of visual information, exclusive focus on COVID-19-related misinformation, and omission of misinformation generated by large language models (LLMs). This dataset contains not only human-generated misinformation but also misinformation produced by LLMs such as ChatGPT, covering 15 disease categories, with data sourced from news articles and tweets. During the dataset construction process, text and their associated images were crawled to ensure the multimodal nature of the data. With broad application scenarios, this dataset is designed to enhance the detection capability of medical misinformation, particularly in distinguishing between human-generated and LLM-generated misinformation.

- 1Med-MMHL: A Multi-Modal Dataset for Detecting Human- and LLM-Generated Misinformation in the Medical Domain弗吉尼亚理工大学 · 2023年



