NeuML/historical-english-books
收藏资源简介:
该数据集是1919年之前的历史英语书籍集合,整合了来自以下四个数据集的书籍:British Library Books、Biodiversity Heritage Library、Library of Congress Selected Digitized Books和Project Gutenberg PG19。根据以下规则进行混合选择:语言为英语;对Biodiversity Heritage Library限制为最受欢迎的自然科学出版物(版本数大于1);对Library of Congress只选择科学主题,因为其他来源已提供良好的普通文学覆盖。数据字段包括书籍ID、来源集合、出版年份、书名、由LLM生成的19世纪散文风格摘要、应用的书域类别标签、OCR质量评分(0.0到1.0)以及书籍全文。
This dataset is a collection of historical English books from before 1919. It aggregates books from the following datasets into a single dataset: British Library Books, Biodiversity Heritage Library, Library of Congress Selected Digitized Books, and Project Gutenberg PG19. A blend of each of the datasets was selected using the following rules: Language = English; limit natural science to most popular publications (>1 edition) from Biodiversity Heritage Library; select only science subjects from Library of Congress collection given good general literature coverage from other sources. Fields include id, source, year, title, abstract (LLM-generated abstract written in prose of 1800s), label (domain category applied to book), quality (OCR quality from 0.0 to 1.0), and text (full text of the book).




