遇见数据集

Multilingual Segmentation Dataset for Historical Prose (13th–16th c.)

收藏
Zenodo2025-08-29 更新2026-05-29 收录
官方服务:

资源简介:

This dataset was developed to train a multilingual sentence segmentation model, used as a pre-processing step in the automatic alignment of historical texts with Aquilign, a multilingual alignment tool developed by our team. The corpus provides training material for sentence-level segmentation in historical prose from the 13th to 16th centuries. Texts were selected for their genre diversity (narrative, didactic, legal, theological, scholarly prose) and for their ability to reflect editorial, orthographic, and linguistic variation across time, geography, and scribal practices. The current version of the corpus (v1) includes approximately 50,000 segmented excerpts across seven historical languages (Latin, French, Castilian, Catalan, Portuguese, Italian, and English). Segment boundaries are annotated using the pound sign (£), typically corresponding to sentences or syntactic units. The corpus does not include part-of-speech tagging or syntactic annotation — only sentence-level segmentation.

提供机构:
Zenodo
创建时间:
2025-08-29
二维码
社区交流群
二维码
科研交流群
商业服务