Japanese-English Subtitle Corpus (JESC)
收藏资源简介:
Japanese-English Subtitle Corpus (JESC)是由斯坦福大学等机构创建的一个大型日英平行语料库,专注于对话式对话这一未被充分代表的领域。该数据集包含超过320万条日英平行句对,是目前最大的公开可用数据集之一。JESC通过网络爬取和自动对齐的电视剧和电影字幕构建而成,其创建过程包括多种新颖的预处理步骤,以确保高单语流畅性和准确的跨语言对齐。该数据集主要用于解决日英语言对在机器翻译中的资源稀缺问题,特别是在非正式对话领域。
Japanese-English Subtitle Corpus (JESC) is a large-scale Japanese-English parallel corpus created by Stanford University and other institutions, focusing on the underrepresented conversational dialogue domain. This dataset contains over 3.2 million Japanese-English parallel sentence pairs, ranking among one of the largest publicly available corpora to date. JESC is constructed from TV drama and movie subtitles crawled from the web and automatically aligned, and its development pipeline incorporates several innovative preprocessing steps to ensure high monolingual fluency and accurate cross-language alignment. This corpus is primarily designed to address the resource scarcity issue of Japanese-English language pairs in machine translation, especially in the informal conversational domain.

- 1JESC: Japanese-English Subtitle Corpus斯坦福大学 · 2018年



