wmt24pp-ce
收藏资源简介:
该数据集是WMT24++基准中车臣语(Chechen)部分的参考译文。它基于原始WMT24++基准(涵盖55种语言)的俄语版本,由专业母语车臣语译者人工翻译而成。数据集在翻译过程中严格遵循了西里尔字母Palochka的大小写语法规则。数据集规模小于1000个样本,适用于机器翻译模型的训练与评估,特别关注车臣语与俄语间的翻译质量。翻译过程中使用的详细指南将在后续论文中发布。
This dataset is the reference translation of the Chechen portion of the WMT24++ benchmark. It is based on the Russian version of the original WMT24++ benchmark (covering 55 languages), and was manually translated by professional native Chechen translators. The dataset strictly follows the case grammar rules of the Cyrillic letter Palochka during translation. The dataset size is less than 1000 samples, suitable for training and evaluation of machine translation models, with a particular focus on translation quality between Chechen and Russian. Detailed guidelines used in the translation process will be released in a subsequent paper.
WMT24++车臣语参考翻译数据集
本数据集为WMT24++基准测试的车臣语(ce)参考翻译版本,由专业母语译者基于俄语版本人工翻译而成。
基本信息
- 许可证:Apache 2.0
- 任务类型:翻译
- 语言:车臣语(ce)
- 数据规模:少于1K条样本
内容说明
- 数据集包含车臣语参考译文,对应原始WMT24++基准测试(覆盖55种语言)中的俄语源文本。
- 翻译过程中严格遵循车臣语语法规范,在语法正确的前提下同时使用西里尔字母“Palochka”(Ӏ)的大小写两种形式。
- 原始WMT24++基准数据集可参考:https://huggingface.co/datasets/google/wmt24pp
翻译与质量控制
- 翻译工作由一名专业车臣语母语译者完成。
- 翻译所用指南将随数据集相关论文一同发布。




