文言文(古文)- 现代文平行语料
收藏资源简介:
这是一个非常全的文言文(古文)- 现代文平行语料,基本涵盖了大部分经典古籍著作。从文学角度出发,本项目将所有古文原文整理至文件夹 `古文原文` 中,并对每本古籍,按篇章/章节进行划分与展示,正文部分存于各章节下的 `text.txt` 中,例如 `论语/学而篇/text.txt` ,`孟子/梁惠王章句上/第一节/text.txt` 。对于平行数据,本项目整理至文件夹 `双语数据` 中,这些双语数据是以句子级别为单位进行划分,本项目提供了原文、译文、双语三种数据格式,例如:`论语/学而篇/source.txt` 、 `论语/学而篇/target.txt` 、 `论语/学而篇/bitext.txt` 。注:所有数据均按行保留了古文原文的相对顺序,即数据非打乱。
This is a comprehensive parallel corpus of classical Chinese (ancient texts) and modern Chinese, covering most of the classic ancient books. From a literary perspective, this project organizes all the original ancient texts into the folder `古文原文`, and for each ancient book, it is divided and displayed by chapters/sections, with the main text stored in `text.txt` under each section, such as `论语/学而篇/text.txt`, `孟子/梁惠王章句上/第一节/text.txt`. For the parallel data, this project organizes it into the folder `双语数据`, where the bilingual data is divided at the sentence level. The project provides three data formats: original text, translated text, and bilingual text, such as `论语/学而篇/source.txt`, `论语/学而篇/target.txt`, `论语/学而篇/bitext.txt`. Note: All data retains the relative order of the original ancient texts by line, meaning the data is not shuffled.
文言文(古文)- 现代文平行语料概述
数据集结构
- 古文原文:包含327本书籍,按篇章/章节划分,正文存于各章节下的
text.txt文件中。 - 双语数据:包含97本书籍,提供原文、译文、双语三种数据格式,以句子级别对齐,共计972467个句对。
数据特点
- 数据来源于互联网,经过处理后形成句子级别对齐的双语数据。
- 采用归一化编辑距离算法与长度比指标进行核心对齐。
- 双语数据文件夹中的古文数据量少于古文原文文件夹,因部分古文无译文或译文残缺。
统计信息
- 古文原文包含327本书籍。
- 双语数据包含97本书籍,共计972467个句对。
数据来源与声明
- 所有数据均注明出处,详见各书目下的
数据来源.txt文件。 - 原始数据的最终解释权归相关数据来源方所有。
更新历史
- v2.0(2023年3月):重新整理数据,保留详尽的原始数据信息,并注明出处。
- v1.0(2022年2月):数据的初始整理。




