TwiBot-22
收藏资源简介:
该数据集名为TwiBot-22,是一个基于图的全面的推特机器人检测基准,它提供了迄今为止最大的数据集,涵盖了推特网络中多样化的实体和关系,其标注质量显著优于现有数据集。该数据集采用了弱监督学习策略来生成高质量的标签,并通过多种采样策略确保了不同类型用户的代表性。数据集的规模包括9293万节点和1.7019亿条边,其任务是进行推特机器人检测。
The dataset named TwiBot-22 is a comprehensive graph-based benchmark for Twitter bot detection. It features the largest-scale dataset to date, covering diverse entities and relationships within Twitter networks, and its annotation quality markedly outperforms existing datasets. This dataset employs a weak supervision learning strategy to generate high-quality labels, and adopts multiple sampling strategies to ensure the representativeness of different user categories. With a scale of 92.93 million nodes and 170.19 million edges, the core task of this benchmark is Twitter bot detection.




