遇见数据集

PST-Bench

收藏
arXiv2024-02-25 更新2024-06-21 收录
官方服务:

资源简介:

PST-Bench是由清华大学计算机科学与技术系和Zhipu AI联合创建的专业标注数据集,专注于计算机科学领域。该数据集包含1576篇论文及其55,014条相关引用,旨在通过高精度和不断增长的数据支持自动算法的发展,以扩展科学知识的进化理解。数据集的创建过程涉及专业的在线论文阅读组,确保高质量和持续的专业标注。PST-Bench的应用领域包括理解科学进化、研究自动论文来源追踪和评估论文影响,旨在通过类比挖掘和思考最终促进创新。

PST-Bench is a professionally annotated dataset jointly created by the Department of Computer Science and Technology of Tsinghua University and Zhipu AI, focusing on the field of computer science. This dataset comprises 1,576 academic papers and 55,014 associated citations, aiming to support the development of automated algorithms with high-precision and continuously growing data to expand the evolutionary understanding of scientific knowledge. The construction of the dataset involves a professional online paper reading group to ensure high-quality and sustained professional annotation. Application scenarios of PST-Bench include understanding scientific evolution, researching automated paper source tracking and evaluating paper impact, with the goal of ultimately promoting innovation through analogical mining and reasoning.

创建时间:
2024-02-25
搜集汇总
数据集介绍
PST-Bench 数据集图片
背景与挑战
背景概述
PST-Bench是一个用于追踪和基准测试出版物来源的数据集,包含通过Grobid API从PDF生成的论文XML文件。该数据集与KDD Cup 2024竞赛相关,支持多种基线方法(如随机森林和SciBERT)进行来源追踪任务,旨在评估学术出版物来源的识别性能。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务