FLORES+ Wu
收藏资源简介:
FLORES+ Wu数据集是由华东师范大学等机构创建的,旨在为吴语机器翻译模型提供训练和评估基准。该数据集包含997条句子,内容直接从英语翻译成吴语,特别是崇明方言。数据集的创建过程包括手动翻译、验证和标准化处理,确保数据的质量和一致性。该数据集主要应用于吴语机器翻译模型的开发和评估,旨在解决吴语这种资源匮乏语言的机器翻译难题。
The FLORES+ Wu dataset was created by institutions including East China Normal University, aiming to provide training and evaluation benchmarks for machine translation models targeting Wu dialect. This dataset includes 997 sentences directly translated from English into Wu dialect, particularly the Chongming dialect. The dataset creation process involves manual translation, validation and standardization procedures to ensure data quality and consistency. It is primarily applied to the development and evaluation of Wu dialect machine translation models, aiming to address the machine translation challenges of under-resourced languages like Wu dialect.

- 1Machine Translation Evaluation Benchmark for Wu Chinese: Workflow and Analysis华东师范大学 · 2024年



