MTMask6M
收藏资源简介:
MTMask6M数据集是由美团、上海交通大学和北京理工大学联合构建的大规模图像-文本掩膜对数据集,包含600万条图像-文本掩膜对。该数据集通过一个创新的图像-文本对齐方法VQAMask构建,旨在为视觉文档理解任务提供空间感知的特征表示学习。数据集的构建过程中,使用了聚类算法对文本区域进行二值化处理,并通过掩膜生成模块来确保图像中的视觉文本与其对应的图像区域在空间上对齐。该数据集的应用领域主要是为了解决视觉文档理解中的空间对齐问题,提升模型在文档理解任务上的性能。
The MTMask6M dataset is a large-scale image-text mask pair dataset jointly constructed by Meituan, Shanghai Jiao Tong University, and Beijing Institute of Technology, which contains 6 million image-text mask pairs. Constructed via an innovative image-text alignment method named VQAMask, this dataset aims to provide spatially-aware feature representation learning for visual document understanding tasks. During its construction process, clustering algorithms are used to binarize text regions, and a mask generation module is employed to ensure the spatial alignment between visual text in the image and its corresponding image region. The main application scenario of this dataset is to address the spatial alignment issue in visual document understanding, thereby improving the performance of models on document understanding tasks.




