Institutional Books 1.0
收藏资源简介:
Institutional Books 1.0 是一个由哈佛图书馆的藏品组成的大型数据集,包含了2420亿个token的公共领域书籍。该数据集由哈佛法学院图书馆、图书馆创新实验室和哈佛图书馆合作创建,旨在为大型语言模型(LLMs)提供高质量的训练数据。数据集涵盖了1,075,899卷,使用250多种语言编写,总共有约2500亿个token。其中,983,004卷被认为是公共领域的,包括原始和后处理的OCR提取文本以及书目、来源和生成的元数据。数据集涵盖了多个世纪,其中60%的内容是在1820年至1920年间出版的。该数据集可用于文学、法律、哲学、科学等领域的训练,旨在解决大型语言模型训练数据稀缺的问题。
Institutional Books 1.0 is a large-scale dataset constructed from the holdings of Harvard Library, which contains public-domain books totaling 242 billion tokens. This dataset was jointly developed by the Harvard Law School Library, the Library Innovation Lab, and Harvard Library, with the objective of supplying high-quality training data for Large Language Models (LLMs). It encompasses 1,075,899 volumes written in more than 250 languages, with an overall token count of approximately 250 billion. Among these volumes, 983,004 are categorized as public domain, including raw and post-processed OCR-extracted text, as well as bibliographic, provenance, and generated metadata. Spanning multiple centuries, 60% of the dataset's content was published between 1820 and 1920. This dataset can be utilized for training in disciplines such as literature, law, philosophy, and science, and aims to resolve the shortage of training data for Large Language Models.




