American Stories
收藏资源简介:
American Stories数据集是由哈佛大学等机构的研究人员开发,包含从美国国会图书馆的公共领域Chronicling America收藏中提取的近2亿扫描图像的结构化文本数据。该数据集覆盖了所有州,内容可追溯至17世纪,主要集中在20世纪初。数据集提供了高质量的数据,可用于预训练大型语言模型,以更好地理解历史英语和历史世界知识。此外,结构化的文章文本便于使用基于变换器的方法进行社会科学应用,如主题分类、内容复制检测和新闻故事聚类。数据集还提供了大规模的银质量数据,用于创新多模态布局分析模型和其他多模态应用。
The American Stories dataset was developed by researchers from Harvard University and other institutions. It contains nearly 200 million structured text data extracted from scanned images in the public-domain Chronicling America collection of the Library of Congress. The dataset covers all U.S. states, with content dating back to the 17th century and primarily focused on the early 20th century. It provides high-quality data suitable for pre-training large language models (LLMs) to better understand historical English and historical world knowledge. Furthermore, its structured article text facilitates the use of Transformer-based methods for social science applications, including topic classification, content copy detection, and news story clustering. Additionally, the dataset offers large-scale silver-standard data for developing innovative multimodal layout analysis models and other multimodal applications.

- 1American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers哈佛大学 · 2023年



