vintage-v2
收藏资源简介:
Vintage-v1 是一个历史书籍文本数据集,旨在为历史语言研究和文本分析提供资源。该数据集整合了来自四个著名历史语料库的书籍:CLMET v3.1(包含 335 本书)、ECCO(3,102 本书)、EEBO(34,937 本书)以及 EVANS(5,013 本书),书籍总数超过 4.3 万本。所有书籍文本均经过清理处理(尽最大努力),从原始的 XML 格式转换为更易读的 Markdown 格式。此外,对于源数据中已标注对话的书籍,其对话内容被专门提取并存储为 JSON 格式,便于结构化访问。数据集整体规模在 1 万到 10 万条之间(以书籍为单位),适用于历史语言学分析、文本风格迁移、对话系统训练或作为大语言模型预训练的历史领域数据。
Vintage-v1 is a historical book text dataset designed to provide resources for historical language research and text analysis. The dataset integrates books from four well-known historical corpora: CLMET v3.1 (containing 335 books), ECCO (3,102 books), EEBO (34,937 books), and EVANS (5,013 books), with a total of over 43,000 books. All book texts have been cleaned (to the best effort) and converted from the original XML format to a more readable Markdown format. Additionally, for books with annotated dialogue in the source data, the dialogue content has been specifically extracted and stored in JSON format for structured access. The overall dataset size ranges from 10,000 to 100,000 entries (in terms of books) and is suitable for historical linguistic analysis, text style transfer, dialogue system training, or as historical domain data for large language model pre-training.
数据集名称
Vintage-v2
数据集描述
Vintage 数据集是一个包含历史书籍文本的集合,所有书籍均经过清洗处理,从 XML 格式转换为 Markdown 格式。其中对话部分已进行标注并提取为 JSON 格式。
数据集规模
- 样本数量:10,000 < n < 100,000
语言
- 英语(en)
标签
- vintage(复古)
- historical(历史)
数据来源
数据集包含以下四个来源的书籍:
- CLMET v3.1:335 本书
- ECCO:3,102 本书
- EEBO:34,937 本书
- EVANS:5,013 本书
许可证
- Creative Commons(CC)
配置文件
- 配置名称:default
- 数据文件:训练集位于
data/train-*.parquet
引用信息
如使用本数据集,请引用: bibtex @misc{vintage-v2, title = {Vintage v2}, author = {Cristi Constantin}, month = {July}, year = {2026}, url = {https://huggingface.co/datasets/croqaz/vintage-v2} }




