M3T - 多模态文档级机器翻译新基准数据集
收藏资源简介:
M3T是一个多模态文档级机器翻译基准数据集,由亚马逊联合马里兰大学和奈良科学技术研究所创建,旨在评估神经机器翻译(NMT)系统在翻译半结构化文档时的性能。该数据集专注于半结构化文档的翻译任务,特别针对PDF文档的视觉复杂性进行设计,以挑战并提升NMT系统在处理真实世界文档时的表现。M3T数据集包含从EUR-Lex、DocLayNet和RVL-CDIP等多个公共数据源收集的文档,覆盖法律、金融等多个领域。文档经过专业翻译和后期编辑,确保翻译质量与原文长度保持在±10%的范围内,以评估系统在保持原文布局方面的能力。该数据集的发布推动了多模态机器翻译技术的发展,解决现有NMT系统在翻译具有复杂布局的文档时的挑战。通过M3T,研究人员可以评估和改进模型在利用视觉线索进行高质量翻译方面的能力。
M3T is a multimodal document-level machine translation benchmark dataset, jointly created by Amazon, the University of Maryland, and the Nara Institute of Science and Technology. It is designed to evaluate the performance of neural machine translation (NMT) systems in translating semi-structured documents. The dataset focuses on the translation tasks of semi-structured documents, specifically designed to challenge and enhance the performance of NMT systems in handling real-world documents with the visual complexity of PDFs. The M3T dataset includes documents collected from various public sources such as EUR-Lex, DocLayNet, and RVL-CDIP, covering multiple fields including law and finance. The documents have been professionally translated and post-edited to ensure that the translation quality and the length of the translations are within ±10% of the original texts, assessing the system's ability to maintain the original layout. The release of this dataset has advanced the development of multimodal machine translation technology, addressing the challenges faced by existing NMT systems in translating documents with complex layouts. Through M3T, researchers can evaluate and improve the models' capabilities in utilizing visual cues for high-quality translations.
数据集概述
数据集名称
M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation
数据集目的
该数据集旨在评估神经机器翻译(NMT)系统在翻译半结构化文档方面的能力,特别关注文档布局等视觉线索对翻译任务的影响。
数据集特点
- 针对文档级NMT系统,考虑了文档布局等视觉元素的重要性。
- 旨在解决现有NMT系统在处理复杂文本布局时的不足。
数据集使用许可
本数据集根据CC-BY-4.0许可证授权。
引用信息
若使用本数据集,请考虑引用以下文献:
@misc{hsu2024m3t, title={M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation}, author={Benjamin Hsu and Xiaoyu Liu and Huayang Li and Yoshinari Fujinuma and Maria Nadejde and Xing Niu and Yair Kittenplon and Ron Litman and Raghavendra Pappagari}, year={2024}, eprint={2406.08255}, archivePrefix={arXiv}, primaryClass={cs.CL} }
同时,也请引用原始数据集:
@inproceedings{pfitzmann-et-al, author = {Pfitzmann, Birgit and Auer, Christoph and Dolfi, Michele and Nassar, Ahmed S. and Staar, Peter}, title = {DocLayNet: A Large Human-Annotated Dataset for Document-Layout Segmentation}, year = {2022}, isbn = {9781450393850}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3534678.3539043}, doi = {10.1145/3534678.3539043}, }
@inproceedings{harley-et-al, author={Harley, Adam W. and Ufkes, Alex and Derpanis, Konstantinos G.}, booktitle={2015 13th International Conference on Document Analysis and Recognition (ICDAR)}, title={Evaluation of deep convolutional nets for document image classification and retrieval}, year={2015}, volume={}, number={}, pages={991-995}, doi={10.1109/ICDAR.2015.7333910} }




