mqud
收藏资源简介:
MQUD(Multimodal Questions Under Discussion)是一个包含1,250个基于科学论文中图形的好奇问题的多模态数据集。每个示例包含一个科学图形、论文上下文、一个问题、一个提取性答案、问题类型以及面向作者的元数据。数据集适用于视觉问答任务,支持多模态(文本和图像)处理。数据规模为1K<n<10K,包含56篇论文的245个图形条目。问题类型分为六类:原因(cause)、比较(comparison)、概念(concept)、结果(consequence)、程度(extent)和程序性(procedural),难度分为“困难”(hard)和“中等”(medium)。数据集提供了JSONL和Parquet格式的文件,以及Hugging Face ImageFolder兼容的行。文本字段经过轻度规范化以提高可读性。数据集当前使用`license: other`作为保守的许可证占位符,建议在公开前确认具体的发布许可证。
MQUD (Multimodal Questions Under Discussion) is a multimodal dataset containing 1,250 curiosity-driven questions based on figures from scientific papers. Each example includes a scientific figure, paper context, a question, an extractive answer, question type, and author-facing metadata. The dataset is suitable for visual question answering tasks and supports multimodal (text and image) processing. The dataset size is 1K<n<10K, comprising 245 figures from 56 papers. Question types are categorized into six classes: cause, comparison, concept, consequence, extent, and procedural, with difficulty levels of hard and medium. The dataset provides files in JSONL and Parquet formats, as well as Hugging Face ImageFolder-compatible rows. Text fields are lightly normalized for readability. The dataset currently uses `license: other` as a conservative license placeholder, and it is recommended to confirm the specific release license before public use.
MQUD: Multimodal Questions Under Discussion 数据集概述
基本信息
- 数据集名称:MQUD(Multimodal Questions Under Discussion)
- 数据集大小:1,250 条数据(1K < n < 10K)
- 语言:英文
- 任务类别:视觉问答(Visual Question Answering)
- 许可证:other(需在公开前确认正式许可)
数据集内容
该数据集包含来自科学论文的 1,250 个基于图表的探究性问题。每个样本包含:
- 科学图像
- 论文上下文
- 问题与抽取式答案
- 问题类型
- 面向作者的元数据
数据字段(公共 JSONL/Parquet 格式)
| 字段 | 说明 |
|---|---|
id、paper_id |
稳定的样本和论文标识符 |
image |
图像路径(位于 imagefolder/images/ 下) |
question、answer |
探究性问题及抽取式答案 |
question_type |
问题类型(共6类) |
difficulty |
标注难度标签 |
paper_title、paper_abstract、figure_caption |
来源论文/图表上下文 |
source_text |
支撑性论文上下文(单文本字段) |
source_paragraphs |
支撑性上下文(段落列表形式) |
数据分布
问题类型分布
| 问题类型 | 数量 |
|---|---|
| cause(原因) | 295 |
| comparison(比较) | 233 |
| concept(概念) | 160 |
| consequence(结果) | 192 |
| extent(程度) | 227 |
| procedural(程序性) | 143 |
难度分布
- hard:361 条
- medium:889 条
统计数量
- 样本总数:1,250
- 来源论文数:56
- 图表条目数:245
- 原始图像路径数:244(归一化后为243)
数据文件
数据集包含以下文件:
data/mqud.jsonl:每行一个 MQUD 问题data/mqud.parquet:Parquet 格式的相同数据imagefolder/metadata.jsonl:兼容 Hugging Face ImageFolder 的行数据metadata/papers.jsonl:每篇源论文的聚合计数metadata/dataset_summary.json:数据集级别的统计metadata/imagefolder_manifest.csv:每个样本的图像副本状态imagefolder/images/:存放所有图像文件
加载方式
python from datasets import load_dataset
ds = load_dataset("imagefolder", data_dir="imagefolder", split="train")
相关论文




