Heterogeneous Text-Attributed Graph Datasets
收藏资源简介:
本文介绍了一系列多尺度异构文本属性图(HTAG)数据集,这些数据集涵盖了电影、社区问答、学术、文学和专利等多个领域。数据集规模从小型(24K节点,104K边)到大型(5.6M节点,29.8M边)不等,提供了丰富的文本内容和多样的关系类型。数据集的创建过程包括从多个公开数据源收集元数据,并通过预训练语言模型(PLMs)生成文本特征。这些数据集旨在支持机器学习模型在异构文本属性图上的真实和可重复评估,特别适用于图神经网络(GNNs)的研究和应用,以解决复杂网络中的节点分类和关系预测等问题。
This paper introduces a series of multi-scale heterogeneous text-attributed graph (HTAG) datasets, which cover multiple domains including movies, community question answering, academia, literature, and patents. These datasets range in scale from small (24K nodes, 104K edges) to large (5.6M nodes, 29.8M edges), providing rich textual content and diverse relationship types. The construction of these datasets involves collecting metadata from multiple public data sources and generating textual features via pre-trained language models (PLMs). These datasets are designed to support authentic and reproducible evaluations of machine learning models on heterogeneous text-attributed graphs, and are particularly suitable for the research and application of graph neural networks (GNNs) to solve tasks such as node classification and relation prediction in complex networks.



