遇见数据集

credi-net/CrediBench

收藏
Hugging Face2026-05-25 更新2026-03-29 收录
官方服务:

资源简介:

CrediBench 1.1是一个大规模、时序性的网络图数据集,专为在线错误信息检测研究设计。数据来源于Common Crawl,并补充了网络爬取和多个可信度信号数据集。数据集包含每月的大规模网络切片,每个月的网络图拥有超过10亿条边和超过4500万个节点,其中节点代表网站域名(如google.com),边代表有向超链接关系(例如,从cbc.ca到reuters.com的边表示cbc.ca网站上的页面包含指向reuters.com页面的超链接)。此外,数据集还提供了文本属性(部分来自Common Crawl和网络爬取),这些文本特征在错误信息检测中起重要作用,并补充了来自Lin等人的可信度评分,以支持监督和半监督学习。数据集由Complex Data Lab @ Mila - Quebec AI Institute、牛津大学、麦吉尔大学、康考迪亚大学、加州大学伯克利分校、蒙特利尔大学和AITHYRA的合作团队策划,采用CC-BY-4.0许可证。统计数据包括2024年10月至2025年5月各月的节点数、边数、最小度、平均度、最大度、叶子节点和边密度等。数据集还包含基于LLM的领域级内容嵌入,用于下游图神经网络模型的特征初始化,并提供了内容比例表格,显示在回归和二元分类任务中,各月数据集的可信与非可信类别的标签数量和内容占比。

CrediBench is a large-scale, temporal webgraph dataset derived from Common Crawl, supplemented with web scraping and multiple datasets for credibility signals. It consists of monthly slices of large-scale web networks, each containing over 1 billion edges and over 45 million nodes, where nodes represent website domains and edges represent directed hyperlink relations. The dataset is enhanced with text attributes (partly from Common Crawl and web scraping) and credibility scores from Lin et al. to facilitate supervised and semi-supervised learning, particularly for misinformation detection. Curated by a collaborative team from the Complex Data Lab @ Mila - Quebec AI Institute, University of Oxford, McGill University, Concordia University, UC Berkeley, University of Montreal, and AITHYRA, it is licensed under CC-BY-4.0. Statistics include node count, edge count, minimum degree, mean degree, maximum degree, leaves, and edge density for months from October 2024 to May 2025. It also includes domain-level content embeddings generated using LLM-based models for feature initialization in GNNs, along with proportion of content tables for regression and binary classification tasks.

提供机构:
credi-net
二维码
社区交流群
二维码
科研交流群
商业服务