arXiv4TGC
收藏资源简介:
arXiv4TGC是一个专为大规模时序图聚类设计的新颖学术数据集,包含arXivAI、arXivCS、arXivMath、arXivPhy和arXivLarge五个子数据集。这些数据集从arXiv开放平台提取,涵盖172个子领域,以论文为节点,引用关系为边,记录了时间依赖的交互。其中最大的数据集arXivLarge包含130万个标记节点和1000万条时序边。数据集创建过程中,首先从原始数据提取节点交互信息,根据领域分类提取边,重编号节点并更新标签列表,最后提供基于位置编码的节点特征。arXiv4TGC不仅适用于时序图聚类,还可用于其他图学习任务,如节点分类,旨在解决现有时序图数据集规模小、标签不可靠的问题。
arXiv4TGC is a novel academic dataset specifically designed for large-scale temporal graph clustering, which includes five sub-datasets: arXivAI, arXivCS, arXivMath, arXivPhy, and arXivLarge. These datasets are extracted from the arXiv open platform, covering 172 sub-fields, with papers as nodes and citation relationships as edges to record time-dependent interactions. The largest sub-dataset, arXivLarge, contains 1.3 million labeled nodes and 10 million temporal edges. During the dataset construction process, node interaction information is first extracted from raw data, edges are extracted based on domain classifications, nodes are renumbered and label lists are updated, and finally node features based on positional encoding are provided. arXiv4TGC is applicable not only to temporal graph clustering but also to other graph learning tasks such as node classification, aiming to address the issues of small scale and unreliable labels of existing temporal graph datasets.




