Novelty Detection Datasets
收藏资源简介:
该数据集是为科学创新性检测(ND)任务量身定制的基准数据集,具有拓扑闭包性和紧凑性。数据集包括营销和自然语言处理(NLP)两个领域的子集,其中营销领域数据集包含470篇研究文章,NLP领域数据集包含3,533篇论文。数据集构建过程利用大型语言模型(LLM)提取和总结每篇论文的核心贡献、假设和方法论,以提高数据集的紧凑性。数据集旨在解决现有ND方法在资源密集和主观性方面的限制,通过LLM的知识蒸馏框架训练思想检索器,以捕获概念而非文本相似度,从而有效地检测研究思想的新颖性。
This dataset is a tailored benchmark for the Scientific Novelty Detection (ND) task, featuring topological closure and compactness. It comprises two subsets from the marketing and Natural Language Processing (NLP) domains: the marketing subset contains 470 research articles, while the NLP subset includes 3,533 academic papers. During the dataset construction process, Large Language Models (LLMs) were utilized to extract and summarize the core contributions, hypotheses, and methodologies of each paper, thereby enhancing the dataset's compactness. This dataset aims to address the limitations of existing ND methods in terms of resource intensiveness and subjectivity. It trains a thought retriever through the knowledge distillation framework of LLMs to capture conceptual rather than textual similarity, thus effectively detecting the novelty of research ideas.

- 1Harnessing Large Language Models for Scientific Novelty Detection南洋理工大学 · 2025年



