遇见数据集

PredCheck: Detecting Predatory Behaviour in Scholarly World

收藏
Mendeley Data2024-03-27 更新2024-06-28 收录
数据链接:
官方服务:

资源简介:

Dataset used in the paper "PredCheck: Detecting Predatory Behaviour in Scholarly World" accepted at JCDL 2020 as a poster. Abstract: High solicitation for publishing a paper in scientific journals has led to the emergence of a large number of open-access predatory publishers. They fail to provide a rigorous peer-review process, thereby diluting the quality of research work and charge high article processing fees. Identification of such publishers has remained a challenge due to the vast diversity of the scholarly publishing ecosystem. Earlier works utilises only the objective features such as metadata. In this work, we aim to explore the possibility of identifying predatory behaviour through text-based features. We propose PredCheck, a four-step classificaton pipeline. The first classifier identifies the subject of the paper using TF-IDF vectors. Based on the subject of the paper, the Doc2Vec embeddings of the text are found. These embeddings are then fed into a Naive Bayes classifier that identifies the text to be predatory or non-predatory. Our pipeline gives a macro accuracy of 95% and an F1-score of 0.89.

本数据集用于2020年国际数字图书馆联合会议(JCDL 2020)以海报形式收录的论文《PredCheck:学术领域掠夺性行为检测》。其论文摘要如下:科研论文发表需求的旺盛催生了大量开放获取型掠夺性出版商(predatory publishers)。此类出版商未提供严格的同行评议流程,不仅稀释了科研成果的质量,还收取高额论文处理费。由于学术出版生态系统的多样性极强,识别此类出版商始终是一项挑战。此前的相关研究仅采用元数据(metadata)这类客观特征。本研究旨在探索基于文本特征识别掠夺性行为的可行性。我们提出了PredCheck这一四步分类流水线:首先利用TF-IDF向量识别论文的主题领域;基于识别出的论文主题,生成文本的Doc2Vec嵌入向量;随后将这些嵌入向量输入朴素贝叶斯分类器(Naive Bayes classifier),以判定文本属于掠夺性还是非掠夺性类别。本分类流水线的宏准确率可达95%,F1值为0.89。

创建时间:
2023-06-28
二维码
社区交流群
二维码
科研交流群
商业服务