CAMMT
收藏资源简介:
CAMMT是一个由人类精心策划的多模态机器翻译基准数据集,包含超过5,800个图像及其英文和区域语言的平行字幕。该数据集覆盖了19种语言和23个地区。通过CAMMT,研究人员评估了五种视觉语言模型在仅文本和文本+图像设置下的翻译质量。结果表明,视觉上下文通常能提高翻译质量,特别是在处理文化特定物品、歧义消解和正确使用性别方面。CAMMT旨在支持构建和评估更具文化细微差别和地区差异的多模态翻译系统。
CAMMT is a human-curated benchmark dataset for multimodal machine translation, containing over 5,800 images paired with parallel subtitles in English and regional languages. This dataset covers 19 languages and 23 regions. Using CAMMT, researchers evaluated the translation quality of five vision-language models under two settings: text-only and text-plus-image. Results show that visual context generally improves translation quality, especially when handling culture-specific items, ambiguity resolution, and proper gender usage. CAMMT aims to support the development and evaluation of multimodal translation systems with greater cultural nuance and regional diversity.
数据集概述
基本信息
- 数据集名称: villacu/cammt
- 下载大小: 689642字节
- 数据集大小: 1379321字节
数据特征
数据集包含以下字段:
ID: 字符串类型regional: 字符串类型English: 字符串类型Conserved_translation: 字符串类型Substituted_translation: 字符串类型Category: 字符串类型Preferred_translation: 字符串类型
数据分块
数据集包含以下分块及其信息:
| 分块名称 | 字节大小 | 样本数量 |
|---|---|---|
| es_mex | 76789 | 323 |
| bn_india | 82705 | 286 |
| om_eth | 51601 | 214 |
| ur_india | 52795 | 220 |
| ig_nga | 36977 | 200 |
| ur_pak | 53728 | 216 |
| zh_ch | 49137 | 308 |
| es_ecu | 84685 | 362 |
| sw_ken | 93951 | 271 |
| kor_sk | 62831 | 290 |
| ru_rus | 52428 | 200 |
| ta_india | 69530 | 213 |
| amh_eth | 55026 | 234 |
| jp_jap | 47779 | 203 |
| fil_phl | 41161 | 203 |
| ms_mys | 61583 | 315 |
| bg_bg | 78441 | 369 |
| es_chl | 55839 | 234 |
| pt_brz | 62512 | 284 |
| ar_egy | 45666 | 203 |
| ind_ind | 40262 | 202 |
| mr_india | 62211 | 202 |
| es_arg | 61684 | 265 |
配置信息
- 配置名称: default
- 数据文件路径: 各分块数据文件路径以
data/{分块名称}-*格式存储

- 1CaMMT: Benchmarking Culturally Aware Multimodal Machine TranslationMBZUAI · 2025年



