遇见数据集

Machine Translation Evaluation Dataset for Amharic

收藏
Zenodo2020-07-31 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

# Machine Translation Evaluation Dataset for Amharic The dataset contains sentences in Amharic and their corresponding translations<br> in English that were collected using crowd sourcing. These ground-truth<br> sentences are from across different domains such as news headlines, social<br> media, Wikipedia and everyday conversation. <br> ## Metadata of files in the dataset amen.tsv<br> - Domain: news | wiki | twitter | convo<br> - Source Sentence: Amharic sentence<br> - Reference Translation: English translation<br> - Google Translate: output of Google Translate<br> - Yandex Translate: output of Yandex Translate <br> enam.tsv<br> - Domain: news | wiki | twitter | convo<br> - Source Sentence: English sentence<br> - Reference Translation: Amharic translation<br> - Google Translate: output of Google Translate<br> - Yandex Translate: output of Yandex Translate <br> ## Reference translations across domains **News**<br> - These are news headlines from Ethiopian news websites. <br> **Wikipedia**<br> - A random sample of sentences from the Amharic Wikipedia. <br> **Twitter**<br> - Amharic Twitter posts on consumer products. <br> **Conversational**<br> - Everyday conversational expressions from Amharic native speakers. <br> ## Evaluation of two systems that provide Amharic translation The dataset also contains evaluation of two commercial systems: [Google<br> Translate](https://translate.google.com/) and [Yandex<br> Translate](https://translate.yandex.com/). Both systems provide free APIs that<br> users can sign up and get access keys. The translations were generated on 14th<br> February 2020.

提供机构:
Zenodo
创建时间:
2020-02-17
二维码
社区交流群
二维码
科研交流群
商业服务