遇见数据集

0601p/Traveling_Namuwiki_Paths

收藏
Hugging Face2026-05-06 更新2026-05-31 收录
官方服务:

资源简介:

Traveling Namuwiki是一个基于Namuwiki页面链接构建的图导航数据集。每个数据示例包含起始页面标题、目标页面标题以及它们之间的一个或多个有效最短路径,路径以中间页面标题列表的形式存储(不包括起始和目标页面)。该数据集旨在用于图机器学习和文本生成任务,特别是模拟在维基百科式知识图中进行导航或路径查找的场景。数据集来源于Hugging Face上的heegyu/namuwiki数据集,并经过处理以包含跳数感知的分层分割,总共有约199.5万行数据,分为训练、验证和测试集。跳数分布从1到10不等,其中大多数路径的跳数在1到4之间。数据模式包括start_title、target_title、paths和hop字段,其中hop表示从起始页面到目标页面的最小跳数,paths是所有实现该最小跳数的最短路径集合。

Traveling Namuwiki is a graph-navigation dataset built from Namuwiki page links. Each example contains a start page, a target page, and one or more valid paths between them, stored as intermediate page-title lists excluding the start and target pages. This dataset is designed for graph machine learning and text generation tasks, specifically simulating navigation or path-finding scenarios in a Wikipedia-style knowledge graph. It was derived from the Hugging Face dataset heegyu/namuwiki and includes hop-aware stratification splits, with approximately 1.995 million rows divided into train, validation, and test sets. The hop distribution ranges from 1 to 10, with most paths having 1 to 4 hops. The schema includes fields for start_title, target_title, paths, and hop, where hop is the minimum number of hops needed to travel from start to target, and paths is the set of all shortest paths achieving that minimum.

提供机构:
0601p
二维码
社区交流群
二维码
科研交流群
商业服务