17th-Century Italian Corpus (1610–1689)
收藏资源简介:
该数据集是由博洛尼亚大学FICLIT数字图书馆构建的17世纪意大利语历史文本语料库,涵盖1610年至1689年间的宗教论文、传记、历史叙事等多种体裁。语料库包含732个对齐的句子对,约15,000个词元,数据来源于原始页面图像经定制OCR处理并辅以GPT-5-mini生成语义规范化版本。创建过程通过提取连续页面范围以保持话语结构,并进行了人工验证以确保语义保真度。该数据集主要用于评估大型语言模型对历史语言的处理能力,旨在探究拼写变异、词汇距离和预训练暴露对模型编码成本与理解难度的影响,为数字图书馆中的语义检索任务提供基准。
This dataset is a 17th-century Italian historical text corpus constructed by the FICLIT Digital Library of the University of Bologna, covering various genres including religious treatises, biographies, historical narratives and more spanning from 1610 to 1689. The corpus contains 732 aligned sentence pairs and approximately 15,000 tokens. Its data is derived from original page images processed via customized OCR, supplemented with semantically normalized versions generated by GPT-5-mini. During the construction process, contiguous page ranges were extracted to preserve discourse structure, and manual validation was conducted to ensure semantic fidelity. This dataset is primarily used to evaluate the processing capabilities of large language models (LLMs) for historical languages, aiming to investigate the impacts of spelling variation, lexical distance and pretraining exposure on model encoding costs and comprehension difficulty, providing a benchmark for semantic retrieval tasks in digital libraries.

- 1How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation博洛尼亚大学 · 2026年



