MAPS数据集
收藏资源简介:
MAPS数据集由莫纳什大学的研究团队创建,旨在评估机器翻译模型在翻译文化元素(如谚语)时的表现。该数据集包含来自英语、德语、孟加拉语、印尼语和汉语的1773条谚语,每条谚语都附有解释、机器翻译和标签,用于区分谚语的比喻意义和字面意义。数据集的创建过程包括从多语言谚语库中扩展数据,并通过人工注释确保翻译的准确性和文化适应性。该数据集的应用领域主要集中在机器翻译研究,特别是跨文化元素的翻译,旨在解决现有翻译模型在处理文化特定表达时的不足。
The MAPS dataset was developed by a research team at Monash University, with the objective of evaluating the performance of machine translation models when translating cultural elements such as proverbs. It comprises 1,773 proverbs from English, German, Bengali, Indonesian and Mandarin Chinese, where each proverb is accompanied by an explanation, a machine translation and labels for distinguishing between its figurative and literal meanings. The dataset construction process involves expanding data from multilingual proverb repositories and ensuring translation accuracy and cultural appropriateness via manual annotation. Its primary application domains focus on machine translation research, particularly the translation of cross-cultural elements, with the goal of addressing the limitations of current translation models when dealing with culture-specific expressions.

- 1Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model莫纳什大学 · 2025年



