遇见数据集

MIZAN

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

自然语言处理中最重要和最重要的任务之一是机器翻译,它现在高度依赖于多语言并行语料库。通过本文,我们介绍了最大的波斯-英语平行语料库,该语料库从文学名著中收集到的句子对超过一百万对。我们还介绍了语料库的获取过程和统计数据,并使用该语料库试验了一个基线统计机器翻译系统。

One of the most critical and fundamental tasks in natural language processing (NLP) is machine translation (MT), which now heavily relies on multilingual parallel corpora. In this work, we introduce the largest Persian-English parallel corpus to date, which contains over one million sentence pairs collected from classic literary works. We also detail the corpus acquisition process and its statistical characteristics, and conduct experiments on a baseline statistical machine translation system using this corpus.

提供机构:
OpenDataLab
创建时间:
2022-05-07
搜集汇总
数据集介绍
MIZAN 数据集图片
背景与挑战
背景概述
MIZAN是一个波斯-英语平行语料库,专门用于机器翻译任务,包含超过一百万句子对,从文学名著中收集而成。该数据集由匹兹堡大学于2018年发布,基于CC BY 4.0许可,是目前最大的波斯-英语平行语料库之一。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务