HEADLINES
收藏资源简介:
HEADLINES数据集是由哈佛大学和国家经济研究局的研究人员创建,包含近4亿条从1920年至1989年的历史英语报纸中提取的语义相似性数据对。该数据集利用了新数字化的美国地方报纸文章,通过深度神经网络方法识别来自同一来源的文章,构建了大规模的语义相似性数据集。HEADLINES数据集不仅规模庞大,而且覆盖了长时间跨度,适用于训练和评估旨在捕捉抽象相似性的模型,如聚类、最近邻检索和语义搜索。此外,该数据集还能用于评估动态语言模型处理持续演变的文本内容的能力,以及大型语言模型处理历史文本的适应性。
The HEADLINES dataset was created by researchers from Harvard University and the National Bureau of Economic Research. It contains nearly 400 million semantic similarity data pairs extracted from historical English newspapers spanning from 1920 to 1989. This dataset leverages newly digitized articles from U.S. local newspapers, employing deep neural network methodologies to identify articles from the same source, thereby constructing a large-scale semantic similarity dataset. Boasting a massive scale and a long temporal coverage, the HEADLINES dataset is suitable for training and evaluating models designed to capture abstract similarity, such as clustering, nearest neighbor retrieval, and semantic search. Furthermore, this dataset can also be used to evaluate the capability of dynamic language models to process continuously evolving textual content, as well as the adaptability of large language models when handling historical texts.

- 1A Massive Scale Semantic Similarity Dataset of Historical English哈佛大学 · 2023年



