Automatically Generated Plagiarism Dataset
收藏资源简介:
该数据集由三个大型语言模型生成,用于检测科学文章中的自动生成文本抄袭,并与其各自的来源进行对齐。数据集包含78,038对文档,分为训练、验证和测试集。数据集的创建过程涉及从arXiv中抽取文档,使用SPECTER模型创建文档嵌入,并基于语义相似度选择最相似的文档。数据集被分为不同类别,以支持对系统性能的详细分析。
This dataset, generated by three large language models, is designed for detecting automatically generated text plagiarism in scientific articles and aligning plagiarized texts with their respective source documents. It contains 78,038 document pairs, which are split into training, validation, and test sets. The dataset creation process involves extracting documents from arXiv, generating document embeddings using the SPECTER model, and selecting the most semantically similar documents based on semantic similarity. The dataset is divided into multiple categories to support detailed analyses of system performance.




