遇见数据集

MS-CXR: Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing

收藏
DataCite Commons2024-11-15 更新2025-04-16 收录
官方服务:

资源简介:

We release a new dataset, MS-CXR, with locally-aligned phrase grounding annotations by board-certified radiologists to facilitate the study of complex semantic modelling in biomedical vision-language processing. The MS-CXR dataset provides 1162 image-sentence pairs of bounding boxes and corresponding phrases, collected across eight different cardiopulmonary radiological findings, with an approximately equal number of pairs for each finding. This dataset complements the existing MIMIC-CXR v.2 dataset and comprises: 1. Reviewed and edited bounding boxes and phrases (1026 pairs of bounding box/sentence); and 2. Manual bounding box labels from scratch (136 pairs of bounding box/sentence). This large, well-balanced phrase grounding benchmark dataset contains carefully curated image regions annotated with descriptions of eight radiology findings, as verified by radiologists. Unlike existing chest X-ray benchmarks, this challenging phrase grounding task evaluates joint, local image-text reasoning while requiring real-world language understanding, e.g. to parse domain-specific location references, complex negations, and bias in reporting style. This data accompany work showing that principled textual semantic modelling can improve contrastive learning in self-supervised vision-language processing.

我们发布了一款全新的MS-CXR数据集,该数据集包含由认证执业放射科医师标注的局部对齐短语接地(locally-aligned phrase grounding)标注,旨在推动生物医学视觉语言处理领域复杂语义建模的研究。MS-CXR数据集包含1162组包含边界框与对应短语的图像-句子对,采集自8种不同的心肺放射学异常表现,且每种异常对应的样本对数量大致均衡。该数据集作为现有MIMIC-CXR v.2数据集的补充,包含两部分内容:1. 经审核与编辑的边界框与短语标注(共1026组边界框-句子对);2. 从零开始手动标注的边界框与短语标注(共136组边界框-句子对)。 这款规模庞大、分布均衡的短语接地基准数据集,包含经放射科医师验证的、针对8种放射学异常表现的精心筛选图像区域标注。与现有胸部X线影像基准数据集不同,该任务具有较强挑战性,可用于评估联合局部图文推理能力,同时需要具备真实场景下的语言理解能力——例如解析特定领域的位置指代、复杂否定句式以及报告撰写风格中的偏差。本数据集配套的相关研究表明,遵循规范的文本语义建模能够提升自监督视觉语言处理中的对比学习性能。

提供机构:
PhysioNet
创建时间:
2024-07-10
搜集汇总
数据集介绍
MS-CXR: Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing 数据集图片
背景与挑战
背景概述
MS-CXR是一个生物医学视觉-语言处理数据集,包含1162个图像-句子对,由放射科医生标注了局部对齐的短语定位,覆盖八种心肺放射学发现,每类发现数量均衡。它旨在通过短语定位任务评估图像-文本联合推理,处理复杂语义如领域特定位置和否定,以改进自监督视觉-语言处理模型。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务