Erya
收藏资源简介:
Erya数据集是由中国人民大学高瓴人工智能学院创建的,目前是最大的古汉语资源,包含88,808,928条古汉语句子和1,941,396,399个字符。该数据集通过从互联网和开放源数据中收集古汉语材料,并经过清洗和分类,形成了包括单语古文数据和古现代平行数据的综合资源。Erya数据集的创建旨在解决古汉语翻译的难题,通过提供丰富的古汉语资源和分类标准,支持古汉语翻译模型的训练和评估,从而促进古汉语文学的现代传播和理解。
The Erya Dataset was developed by the Gaoling School of Artificial Intelligence, Renmin University of China. It currently stands as the largest ancient Chinese language resource, containing 88,808,928 ancient Chinese sentences and 1,941,396,399 characters. This dataset is built by collecting ancient Chinese materials from the Internet and open-source datasets, followed by cleaning and categorization, forming a comprehensive resource that encompasses both monolingual ancient Chinese data and ancient-modern parallel corpora. The development of the Erya Dataset aims to address the challenges inherent in ancient Chinese translation: by providing rich ancient Chinese resources and standardized classification frameworks, it supports the training and evaluation of ancient Chinese translation models, thereby advancing the modern dissemination and comprehension of ancient Chinese literature.

- 1Towards Effective Ancient Chinese Translation: Dataset, Model, and Evaluation高瓴人工智能学院 · 2023年



