MM-Hallu/vqav2-idk
收藏资源简介:
VQAv2-IDK是一个从VQAv2数据集衍生而来的幻觉评估基准数据集。它包含不可回答(即诱导幻觉)的图像-问题对,其中期望的答案是“我不知道”。数据集用于评估视觉问答系统中的幻觉问题,旨在帮助模型识别并避免对无法回答的问题产生错误响应。数据集结构包括训练集(13,807个示例)和验证集(6,624个示例),每个示例包含图像、唯一问题标识符、问题文本、人类提供的答案列表以及指示不可回答性的关键词(如“unknown”、“none”)。
VQAv2-IDK is a hallucination evaluation benchmark derived from the VQAv2 dataset. It consists of unanswerable (hallucination-inducing) image-question pairs where the desired answer is I Dont Know. The dataset is designed to evaluate hallucination in visual question answering systems, helping models recognize and avoid generating incorrect responses for unanswerable questions. It includes a training set (13,807 examples) and a validation set (6,624 examples), with each example containing an image, a unique question identifier, the question text, a list of human-provided answers, and keywords indicating unanswerability (e.g., unknown, none).



