MERIT Dataset
收藏资源简介:
MERIT数据集是由西班牙马德里康普顿斯大学ICAII工程学院技术研究所创建的多模态数据集,专注于学校报告的文本、图像和布局。该数据集包含33,000个样本,适用于视觉丰富的文档理解(VrDU)任务。数据集通过合成生成方法创建,旨在解决数据稀缺和隐私政策问题,同时评估语言模型中的偏见。其应用领域包括视觉语言模型的预训练、语言模型泛化能力的基准测试以及偏见的检测与缓解。
The MERIT dataset is a multimodal dataset developed by the Technical Institute of the ICAII School of Engineering, Complutense University of Madrid, Spain, focusing on text, images and layouts of school reports. It comprises 33,000 samples and is tailored for Visual Rich Document Understanding (VrDU) tasks. Constructed through synthetic generation approaches, this dataset is designed to address data scarcity and privacy policy issues, while evaluating biases in language models. Its application fields include pre-training of vision-language models, benchmarking the generalization capabilities of language models, as well as detection and mitigation of biases.
- 1The MERIT Dataset: Modelling and Efficiently Rendering Interpretable Transcripts西班牙马德里康普顿斯大学ICAII工程学院技术研究所 · 2024年



