PRIM
收藏资源简介:
PRIM数据集是一个用于实际图像多语言机器翻译(IIMMT)的基准数据集。该数据集包含真实世界捕获的单行文本图像,具有复杂背景、多种字体和多样的文本位置,并支持多种语言翻译方向。数据集的构建过程包括从真实世界中收集图像,手动标注目标图像,并使用GPT-4和Google Translate进行多语言翻译。PRIM数据集的创建旨在帮助研究人员更好地模拟现实世界场景,并为IIMMT研究提供更准确的评估标准。该数据集的应用领域包括翻译软件、图像识别和自然语言处理等,旨在解决实际图像中包含的文本的自动翻译问题。
The PRIM dataset is a benchmark dataset for real-world image multilingual machine translation (IIMMT). This dataset includes single-line text images captured from real-world scenarios, which feature complex backgrounds, diverse fonts and varied text positions, and supports multiple language translation directions. The dataset construction process involves collecting images from real-world scenarios, manually annotating the target images, and performing multilingual translation using GPT-4 and Google Translate. The development of the PRIM dataset aims to help researchers better simulate real-world scenarios and provide more accurate evaluation criteria for IIMMT research. Its application fields cover translation software, image recognition, natural language processing and other related areas, with the goal of addressing the automatic translation of text embedded in real-world images.




