MAmmoTH-VL-Instruct-12M
收藏资源简介:
MAmmoTH-VL-Instruct-12M数据集是一个用于视觉指令调优的数据集,包含1200万条数据。该数据集的创建过程包括手动数据源收集、使用多模态大型语言模型(MLLMs)和大型语言模型(LLMs)进行重写,并通过相同的MLLM进行过滤。数据集主要用于数学和科学类别的指令调优,展示了详细的逐步响应。
MAmmoTH-VL-Instruct-12M is a dataset tailored for visual instruction tuning, comprising 12 million data instances. Its development workflow includes three core stages: manual data source collection, rewriting with Multimodal Large Language Models (MLLMs) and Large Language Models (LLMs), and filtering via the same MLLMs used in the rewriting step. This dataset is primarily designed for instruction tuning tasks in the mathematics and science domains, featuring detailed step-by-step responses.
MAmmoTH-VL-Instruct-12M
简介
MAmmoTH-VL-Instruct-12M 是一个简单且可扩展的视觉指令数据重写管道,包含三个步骤:手动数据源收集、使用MLLMs/LLMs进行重写,以及通过相同的MLLM进行过滤。该数据集展示了数学和科学类别中的详细、逐步的响应。
数据分布
MAmmoTH-VL-Instruct (12M) 的数据分布展示了项目的框架。
引用
@article{guo2024mammothvlelicitingmultimodalreasoning, title={MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale}, author={Jarvis Guo and Tuney Zheng and Yuelin Bai and Bo Li and Yubo Wang and King Zhu and Yizhi Li and Graham Neubig and Wenhu Chen and Xiang Yue}, year={2024}, eprint={2412.05237}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2412.05237}, }




