NeuML/historical-english-books-similarity
收藏资源简介:
该数据集是“历史英语书籍”数据集中标题-摘要对和查询-句子对的随机样本。数据主要来自维多利亚时代直至1899年。标题和摘要直接取自上游数据源,每本书按句子分割,每个分割随机选择句子,并确保分割间无重叠。查询使用19世纪散文风格由大型语言模型生成。此外,数据集还包含一个与BEIR兼容的“测试”分割版本,可用于评估基于此数据训练的向量模型的准确性。
This dataset constitutes a random sample of title-summary pairs and query-sentence pairs sourced from the "Historical English Books" dataset. The data primarily covers the Victorian era through 1899. Titles and summaries are directly extracted from upstream data sources. Each book is segmented at the sentence level, with sentences randomly selected from each segment, and no overlap is ensured between these segments. Queries were generated by a large language model in the style of 19th-century English prose. Additionally, the dataset includes a BEIR-compatible "test" split, which can be used to evaluate the accuracy of vector models trained on this dataset.




