dsfsi/afriscience_mt
收藏资源简介:
AfriScience-MT是一个用于科学机器翻译的平行语料库,支持英语和六种非洲语言(阿姆哈拉语、豪萨语、卢干达语、北索托语、约鲁巴语、祖鲁语),涵盖农业、生物化学、生物学、化学、计算机科学、工程学、地理学、健康、本土知识、社会学和统计学等11个科学领域。该数据集由领域专家科学传播者和专业翻译共同开发,包含230篇论文的7,605个英语源句子,每个句子都翻译成六种非洲语言,并在句子和文档级别对齐。此外,数据集还发布了相关论文中的所有模型预测和每轮评估指标,以及共同开发的双语科学术语表,以便复现和扩展基准测试。数据集包括四个配置:语料库(默认)、预测结果、指标和术语表,分别提供训练、开发和测试分割的数据。
AfriScience-MT is a parallel corpus for scientific machine translation, supporting English and six African languages: Amharic, Hausa, Luganda, Northern Sotho, Yoruba, and Zulu. It covers 11 scientific domains including Agriculture, Biochemistry, Biology, Chemistry, Computer Science, Engineering, Geography, Healthcare, Indigenous Knowledge, Sociology, and Statistics. This dataset was co-developed by domain experts, science communicators and professional translators, containing 7,605 English source sentences from 230 academic papers, each translated into the six aforementioned African languages, with alignment at both sentence and document levels. In addition, the dataset also releases all model predictions and per-round evaluation metrics from the associated papers, as well as the co-developed bilingual scientific terminology glossary, to facilitate benchmark reproduction and expansion. The dataset includes four configurations: Corpus (default), Predictions, Metrics, and Terminology Glossary, which respectively provide training, development and test split data.




