遇见数据集

HPAI-BSC/HEART

收藏
Hugging Face2026-07-08 更新2026-06-14 收录
官方服务:

资源简介:

--- license: cc-by-4.0 task_categories: - visual-question-answering language: - en tags: - medical pretty_name: HEART size_categories: - 10K<n<100K dataset_info: features: - name: dataset dtype: string - name: task dtype: string - name: approach dtype: string - name: sample_type dtype: string - name: source_id dtype: string - name: true_class dtype: int64 - name: index dtype: int64 - name: question dtype: string - name: A dtype: string - name: B dtype: string - name: C dtype: string - name: D dtype: string - name: answer dtype: string - name: fake_option dtype: string - name: true_bbox dtype: string - name: true_class_name dtype: string - name: fake_bboxes dtype: string - name: fake_bboxes_classes dtype: string - name: fake_class dtype: float64 - name: fake_class_name dtype: string - name: caption dtype: string - name: syco_trigger dtype: string - name: answer_color dtype: string - name: image_path dtype: image splits: - name: train num_bytes: 1597468148 num_examples: 24311 download_size: 1156244174 dataset_size: 1597468148 configs: - config_name: default data_files: - split: train path: data/train-* --- # HEART <div align="center"> <img src="https://huggingface.co/datasets/HPAI-BSC/HEART/resolve/main/HEART_logo_transparent.png" width="20%" alt="HPAI"/> </div> <hr style="margin: 15px"> <div align="center" style="line-height: 1;"> <a href="https://hpai.bsc.es/" target="_blank" style="margin: 1px;"> <img alt="Web" src="https://img.shields.io/badge/Website-HPAI-8A2BE2" style="display: inline-block; vertical-align: middle;"/> </a> <a href="https://huggingface.co/HPAI-BSC" target="_blank" style="margin: 1px;"> <img alt="Hugging Face" src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-HPAI-ffc107?color=ffc107&logoColor=white" style="display: inline-block; vertical-align: middle;"/> </a> <a href="https://github.com/HPAI-BSC" target="_blank" style="margin: 1px;"> <img alt="GitHub" src="https://img.shields.io/badge/GitHub-HPAI-%23121011.svg?logo=github&logoColor=white" style="display: inline-block; vertical-align: middle;"/> </a> </div> <div align="center" style="line-height: 1;"> <a href="https://www.linkedin.com/company/hpai" target="_blank" style="margin: 1px;"> <img alt="Linkedin" src="https://img.shields.io/badge/Linkedin-HPAI-blue" style="display: inline-block; vertical-align: middle;"/> </a> <a href="https://bsky.app/profile/hpai.bsky.social" target="_blank" style="margin: 1px;"> <img alt="BlueSky" src="https://img.shields.io/badge/Bluesky-HPAI-0285FF?logo=bluesky&logoColor=fff" style="display: inline-block; vertical-align: middle;"/> </a> <a href="https://linktr.ee/hpai_bsc" target="_blank" style="margin: 1px;"> <img alt="LinkTree" src="https://img.shields.io/badge/Linktree-HPAI-43E55E?style=flat&logo=linktree&logoColor=white" style="display: inline-block; vertical-align: middle;"/> </a> </div> <div align="center" style="line-height: 1;"> <a href="https://creativecommons.org/licenses/by/4.0/deed.en" target="_blank" style="margin: 1px;"> <img alt="License" src="https://img.shields.io/badge/license-CC%20BY%204.0-green" style="display: inline-block; vertical-align: middle;"/> </a> </div> ## Table of Contents - [Table of Contents](#table-of-contents) - [Dataset Summary](#dataset-summary) - [Dataset Sources](#dataset-sources) - [Dataset Modalities](#dataset-modalities) - [Dataset Distribution](#dataset-distribution) - [Dataset Creation](#dataset-creation) - [Dataset Structure](#dataset-structure) - [Considerations for Using the Data](#considerations-for-using-the-data) - [Social Impact of Dataset](#social-impact-of-dataset) - [Discussion of Biases](#discussion-of-biases) - [Other Known Limitations](#other-known-limitations) - [Additional Information](#additional-information) - [Licensing Information](#licensing-information) - [Citation Information](#citation-information) ## Dataset Summary The HEART dataset is composed of multiple splits that differ in the type of injected cues. It includes a *baseline* split with no injected cues and four cue-based splits. Each cue is instantiated in two variants: *assistive*, where the cue is consistent with the GT, and *adversarial*, where the cue supports an incorrect option. Cue types, summarized below, fall into two categories: *prompt-only* cues where the image is unchanged, and *overlay* cues where information is rendered onto the image. - **Baseline.** No cue is added to the original question and available options. These are constructed by shuffling the GT option with wrong alternatives to establish baseline performance. - **Sycophancy (prompt-only).** The prompt asserts a suggested answer, testing whether the model follows the assertion over visual evidence. - **Prompt captions (prompt-only).** A caption-like description is included in the text prompt. In adversarial samples it is constructed to support a distractor label or box. - **Image captions (overlay).** The caption is rendered onto the image as a visual overlay, enabling evaluation of visually embedded textual cues under the same assistive/adversarial setup. - **Legends (overlay; detection only).** Candidate boxes are shown with a legend, mapping box identifiers to labels. The adversarial variant modifies the mapping to support a distractor box. <div style="display: grid; grid-template-columns: repeat(2, 1fr); gap: 2px;"> <img src="https://huggingface.co/datasets/HPAI-BSC/HEART/resolve/main/baseline_example.png" /> <img src="https://huggingface.co/datasets/HPAI-BSC/HEART/resolve/main/sycophantic_example.png" /> <img src="https://huggingface.co/datasets/HPAI-BSC/HEART/resolve/main/image_caption_example.png" /> <img src="https://huggingface.co/datasets/HPAI-BSC/HEART/resolve/main/legend_example.png" /> </div> ### Dataset Sources Data sources used to generate the HEART dataset: - **[ARCADE](https://www.nature.com/articles/s41597-023-02871-z)**: a dataset consisting of *X-ray angiography* images of coronary arteries. From the original dataset, we use the test partition for the stenosis detection task (300 images), which contains annotations by medical experts marking regions affected by atherosclerotic plaques. - **[Breast-Lesions-USG](https://www.nature.com/articles/s41597-024-02984-z)**: a dataset containing 252 *breast ultrasound* scans from different patients, manually annotated with benign and malignant lesions. - **[DENTEX](https://zenodo.org/records/7812323)**: a dataset composed of *panoramic dental X-rays* collected from three different institutions, with images annotated by experts. We use the subset of 250 fully labeled X-rays for abnormal tooth detection. These include four specific diagnosis categories: caries, deep caries, periapical lesions, and impacted teeth. - **[BRISC2025](https://arxiv.org/pdf/2506.14318)**: a dataset composed of 1,000 (test partition only) annotated *MRI scans* for brain tumor segmentation and classification, covering glioma, meningioma, pituitary tumors, and non-tumorous cases. - **[HyperKvasir](https://www.nature.com/articles/s41597-020-00622-y)**: a multi-class image dataset for *gastrointestinal endoscopy*. Since the segmented class corresponds only to the polyps category (1,000 images), we selected images from several other pathological finding classes to form the non-polyps category (857 images in total). - **[FractAtlas](https://www.nature.com/articles/s41597-023-02432-4)**: a dataset composed of 4,083 annotated *X-ray* images for fracture detection and localization, covering wrist, ankle, hip, shoulder, and other common fracture sites. Only 1434 images are utilized to keep a balance between fractured and non-fractured samples. - **[PALM](https://www.nature.com/articles/s41597-024-02911-2)**: a dataset composed of 1,200 annotated *fundus photographs* for retinal disease classification covering pathological myopia and normal controls and optic disc segmentation. - **[ISIC 2017](https://arxiv.org/abs/1710.05006)**: a dataset composed of 2,750 annotated *dermoscopic images* of skin lesions, out of which 600 belong to the test partition, each paired with a diagnosis. It covers three diagnostic categories: melanoma (117 images), seborrheic keratosis (90 images), and benign nevi (393 images), and includes expert segmentations. ### Dataset Modalities The image illustrates examples from all medical imaging modalities included in the HEART dataset. <p align="center"> <img src="https://huggingface.co/datasets/HPAI-BSC/HEART/resolve/main/modalities.png" width="60%"> </p> ### Dataset Distribution For each adversarial subset, we generate 250 samples per configuration: 250 for the assistive and 250 for the adversarial setup. Datasets containing only detection tasks (ARCADE and DENTEX) yield a total of 2,250 samples each, distributed among the baseline, sycophancy, legends, image captions, and prompt captions subsets. Datasets including both classification and detection tasks (BRISC2025, HyperKvasir, FractAtlas, ISIC2017, and PALM) produce 4,000 samples each, as subsets are created for both task types. The Breast-Lesions-USG dataset contains fewer samples (3,811) due to the inability to generate valid false bounding boxes for certain images (*e.g.,* real bounding boxes cover the entire image). Overall, the final **HEART** dataset consists of **24,311 samples**. ## Dataset Creation HEART is created with [BAIT](https://huggingface.co/HPAI-BSC/BAIT), which takes as input an existing dataset with images and appropriate annotations (image-level labels for classification and bounding boxes with category labels for detection) and a configuration file specifying question templates, option sampling rules, enabled cue types, and, when applicable, overlay rendering settings. For each selected task and cue type, BAIT outputs multiple-choice samples containing an image reference (original or rendered), a prompt, a set of answer options, the correct option, and metadata identifying the task. The same process is applied to every sample in the source data: 1) Input loading and validation: each input sample is first validated against missing or corrupted files, ensuring all required metadata (e.g., annotations or labels) are present and well-formed. 2) Candidate option construction: for classification, BAIT forms a set containing the GT label and distractors. For detection, it forms a set containing the GT box and incorrect boxes. When insufficient negative boxes are available, BAIT can generate plausible distractor ("fake") boxes (see subsection A.1) that * Do not overlap with ground truth boxes * Are sampled from regions with similar low-level statistics (e.g., mean/std of pixel intensities) * Maintain realistic size variation relative to GT boxes * Avoid predominantly empty regions 3) Sample instantiation and cue injection: BAIT selects a template and fills placeholders using GT information (assistive) or distractor information (adversarial), then it assembles the final multiple-choice prompt. If a cue type requires an in-image overlay (e.g., captions or legends rendered on the image), it also outputs a rendered copy of the image. ## Dataset Structure HEART is composed of 8 subsets, each presenting the same structure. <div> <img src="https://huggingface.co/datasets/HPAI-BSC/HEART/resolve/main/Dataset Structure.svg" /> </div> ## Considerations for Using the Data ### Licesing Information The dataset is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) ### Citation Information ``` @inproceedings{suarez2026heart, title={HEART Attacks: Healthcare Evaluation of Adversarial RobusTness}, author={Su{\'a}rez-Fern{\'a}ndez, Mart{\'\i}n and Lopez-Cuena, Enrique and Guasch-Mart{\'\i}, Jaume and Garcia-Gasulla, Dario and Arias-Duart, Anna}, booktitle={The 2026 ACM Conference on Fairness, Accountability, and Transparency}, pages={321--351}, year={2026} } ```

The HEART dataset is composed of multiple splits that differ in the type of injected cues. It includes a baseline split with no injected cues and four cue-based splits. Each cue is instantiated in two variants: assistive, where the cue is consistent with the GT, and adversarial, where the cue supports an incorrect option. Cue types, summarized below, fall into two categories: prompt-only cues where the image is unchanged, and overlay cues where information is rendered onto the image. - Baseline: No cue is added to the original question and available options. These are constructed by shuffling the GT option with wrong alternatives to establish baseline performance. - Sycophancy (prompt-only): The prompt asserts a suggested answer, testing whether the model follows the assertion over visual evidence. - Prompt captions (prompt-only): A caption-like description is included in the text prompt. In adversarial samples it is constructed to support a distractor label or box. - Image captions (overlay): The caption is rendered onto the image as a visual overlay, enabling evaluation of visually embedded textual cues under the same assistive/adversarial setup. - Legends (overlay; detection only): Candidate boxes are shown with a legend, mapping box identifiers to labels. The adversarial variant modifies the mapping to support a distractor box. The dataset is built from eight source medical imaging datasets covering modalities such as X-ray angiography, breast ultrasound, panoramic dental X-rays, brain MRI, gastrointestinal endoscopy, fracture X-rays, fundus photography, and dermoscopic images, totaling 24,311 samples. It is designed to evaluate the adversarial robustness of medical visual question answering models.

提供机构:
HPAI-BSC
搜集汇总
数据集介绍
HPAI-BSC/HEART 数据集图片
构建方式
HEART数据集由BAIT框架生成,该框架以现有具备图像与注释的数据集为输入,辅以配置文件设定问题模板、选项采样规则及线索类型。处理每个样本时,首先验证输入数据的完整性,随后基于真实标签或检测框构建候选选项集(分类任务使用标签与干扰项,检测任务使用真实框与错误框)。若负样本不足,BAIT可在无重叠且维持尺寸合理性的前提下生成伪框。接着,根据预设线索类型(协助性或对抗性)填充模板,形成多项选择问题,并针对覆盖式线索同步输出经渲染的图像副本。整个过程覆盖8个医学影像源,最终生成24,311个样本。
使用方法
使用HEART数据集时,用户可通过HuggingFace加载默认配置,其中训练集包含24,311个样本。每个样本以多项选择格式呈现,包含问题文本、四个候选选项(标注为A至D)、正确答案及元数据(如任务类型、线索类型、真实与伪框坐标)。研究人员可直接利用这些预构建的对抗性样本评估模型鲁棒性,或通过dataset字段筛选特定子集(如仅使用盲从对抗样本)。数据集遵循CC BY 4.0许可,适合学术与非商业应用,使用者应引用原论文以尊重创作者贡献。
背景与挑战
背景概述
HEART(Healthcare Evaluation of Adversarial RobusTness)数据集由西班牙巴塞罗那超级计算中心(BSC)的HPAI团队于2026年创建,旨在系统评估医学视觉-语言模型在面对对抗性线索时的鲁棒性。该数据集整合了ARCADE(冠状动脉造影)、Breast-Lesions-USG(乳腺超声)、DENTEX(牙科X光)、BRISC2025(脑部MRI)、HyperKvasir(胃肠内镜)、FractAtlas(骨折X光)、PALM(眼底照片)和ISIC2017(皮肤镜图像)等八个公开医学影像数据集,覆盖X光、MRI、超声、内镜、眼底及皮肤镜等多种模态。核心研究问题聚焦于多模态大模型在临床场景中是否易受文本或视觉线索的误导,从而影响诊断准确性。HEART通过注入奉承性断言、图像描述和标签图例等四种线索类型,分别构建辅助性与对抗性样本,为评估模型在复杂临床环境下的决策可靠性提供了标准化基准,对推动可信医疗AI的发展具有重要影响力。
当前挑战
HEART数据集面临的领域挑战在于医学视觉问答任务中模型对非诊断性线索的过度依赖,这可能导致临床误诊风险。具体而言,模型可能忽略图像中的病理证据,转而采纳提示中的误导性断言或视觉叠加信息,暴露出其在对抗性环境下缺乏鲁棒性的弱点。数据集构建过程中也面临多重技术挑战:需从八个异质医学数据源中统一提取并标准化标注格式,针对检测任务中部分图像缺乏足够负样本的情况,开发了启发式算法生成不重叠、符合局部纹理统计的假框;同时需设计细微的视觉叠加(如图例映射篡改)以区分模型是否真正理解内容而非依赖表面相关性。此外,平衡分类与检测任务在各子集中的样本分布,以及确保辅助性与对抗性变体的严格一致性,也构成了工程上的复杂挑战。
常用场景
经典使用场景
在医疗影像分析领域,HEART数据集被广泛用于评估视觉语言模型在医学多模态推理任务中的鲁棒性,尤其是针对视觉问答任务中模型对偏见线索的敏感性。该数据集通过精心设计的辅助性和对抗性提示,如迎合性断言、嵌入文本覆盖和图像图例,系统性地考察模型在面对与真实标签一致或矛盾的误导信息时的决策行为,成为衡量模型是否真正基于图像内容而非虚假关联进行推理的标杆。
解决学术问题
HEART数据集直击当前医学人工智能研究中的关键痛点——模型在面临对抗性干扰时的脆弱性。它解决了如何量化评估视觉语言模型在医疗诊断场景中对特定类型错误信息的抵抗力问题,揭示了模型倾向于依赖提示中的断言而非视觉证据的现象,为构建更可靠的医学AI系统提供了方法论基础。该数据集的出现推动了关于模型可解释性、安全性和公平性的学术讨论,尤其是在高风险医疗决策背景下,强调了模型必须具备对抗虚假线索的稳健性。
实际应用
在实际临床场景中,HEART数据集可辅助开发更可信的辅助诊断系统,帮助筛选那些在放射学报告或患者病史中存在常见误导信息时仍能保持准确性的模型。例如,当影像报告中包含不准确的历史记录或术语时,基于HEART评估的模型能够避免被干扰,从而减少误诊风险。此外,该数据集还可用于训练医学教育中的AI导师系统,使模型能够识别并纠正学习者的常见认知偏差。
数据集最近研究
最新研究方向
在医疗人工智能领域,随着视觉语言模型(VLM)在临床辅助决策中的深入应用,模型对于提示中注入的虚假或误导性线索(如谄媚性断言、假字幕、视觉叠加干扰)的鲁棒性成为前沿焦点。HEART数据集应运而生,它通过系统性构建基线、谄媚、字幕(文本与叠加)及图例等五类线索变体,模拟真实临床环境中可能存在的语义或视觉偏差,开创性地评估VLM在X射线、超声、MRI、眼底照相等多种医学模态下的抗干扰能力。该研究呼应了医疗AI安全性的热点事件——即模型过度依赖文本模式而忽视视觉证据的脆弱性,其意义在于为构建高可信度、可解释的医疗多模态系统提供了标准化测试基准,推动模型在辅助诊断中实现真正的视觉证据优先。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务