遇见数据集

Embedding the Web: An Open Billion-Scale Repository for Knowledge Graph Embeddings

收藏
Zenodo2025-05-08 更新2026-04-07 收录
官方服务:

资源简介:

We present a new resource of knowledge graph embeddings generated from Web Data Commons dataset--the largest collection of structured data from the web. The whole dataset encompasses 97,689,391,384 RDF triples extracted from over 21,968,201 websites, posing a big scalability challenge for embedding algorithms. To address this challenge, we develop a methodology that partitions the massive web-scale knowledge graph by website domain and learns embeddings from each domain-specific subgraph independently using the state-of-the-art knowledge graph embedding model DECAL. This divide-and-conquer approach, coupled with distributed training on high-performance computing clusters, allows us to produce vector representations for approximately 20.9 billion distinct IRIs--surpassing previous embedding resources by more than two orders of magnitude. Furthermore, our approach is follows a never-ending embedding paradigm, continuously updating the dataset to incorporate newly emerging IRIs and domains, thereby ensuring the resource remains comprehensive and up to date. The resulting embeddings, dubbed Whale-embeddings, are publicly available and constitute the largest knowledge graph embedding resource released to date.

创建时间:
2025-05-08
二维码
社区交流群
二维码
科研交流群
商业服务