masakhane/afriscience_mt
收藏资源简介:
AfriScience-MT 是一个平行科学机器翻译语料库,专为英语和六种非洲语言(阿姆哈拉语、豪萨语、卢干达语、北索托语、约鲁巴语、祖鲁语)设计,由专家科学传播者和专业翻译共同开发,涵盖11个科学领域(农业、生物化学、生物学、化学、计算机科学、工程学、地理学、健康、本土知识、社会学、统计学)。该数据集包含230篇论文的7,605个英语源句子,每个句子都翻译成六种非洲语言,并在句子和文档级别对齐。此外,数据集还发布了伴随论文中的所有模型预测和每轮评估指标,支持基准的复现和扩展,无需重新运行实验。构建过程包括两个阶段:领域专家首先生成每篇论文的通俗摘要,然后由专业翻译进行翻译,并共同开发双语科学术语表以填补标准化术语的空白。数据集旨在通过文本翻译促进非洲科学的去殖民化。
A parallel scientific machine-translation corpus for English + six African languages (Amharic, Hausa, Luganda, Northern Sotho, Yorùbá, isiZulu), co-developed with expert science communicators and professional translators across 11 scientific domains (Agriculture, Biochemistry, Biology, Chemistry, Computer Science, Engineering, Geography, Health, Indigenous Knowledge, Sociology, Statistics). Alongside the corpus we release every model prediction and per-run metric behind the accompanying paper, so the benchmark can be reproduced and extended without re-running any experiments. The corpus includes 230 papers and 7,605 English source sentences, each translated into the six African target languages and aligned at both the sentence and document level. Construction involves a two-stage process: domain experts produce lay summaries, and translators render them into target languages while co-developing bilingual scientific glossaries. The dataset aims towards decolonizing science in Africa through text translation.




