WikiLinks
收藏资源简介:
WikiLinks数据集包含满足以下两个条件的网页:a. 包含至少一个指向维基百科的超链接;b. 该超链接的锚文本与目标维基百科页面的标题紧密匹配。每个维基百科页面代表一个实体(或概念或想法),锚文本作为实体的提及。数据集通过遍历Google的网页索引获得,包含约1100万文档、300万实体和4000万提及。数据集分为109个文本文件,每个文件详细记录了URL、提及和标记等信息。
The WikiLinks dataset comprises web pages that meet the following two criteria: a. They contain at least one hyperlink to Wikipedia; b. The anchor text of the hyperlink closely matches the title of the target Wikipedia page. Each Wikipedia page represents an entity (or concept or idea), with the anchor text serving as a mention of the entity. The dataset was obtained by traversing Google's web index and includes approximately 11 million documents, 3 million entities, and 40 million mentions. The dataset is divided into 109 text files, each detailing information such as URLs, mentions, and tags.
数据集概述
数据集名称
WikiLinks
数据集内容
- 包含至少一个指向Wikipedia的超链接的网页。
- 超链接的锚文本与目标Wikipedia页面的标题紧密匹配。
- 每个Wikipedia页面代表一个实体(或概念或想法),锚文本作为实体的提及。
数据集格式
- 数据集分为109个文本文件。
- 每个文件格式包括:
- URL:网页的URL。
- MENTION:提及的实体,包括提及字符串、字节偏移和目标URL。
- TOKEN:页面上的最少10个不频繁令牌,包括令牌字符串和字节偏移。
数据集统计
- 文档数量:1100万
- 实体数量:300万
- 提及数量:4000万
数据集使用许可
- 遵循Creative Commons Attribution 3.0 Unported (CC BY 3.0)许可。
- 允许复制、分发、传输和改编,以及商业使用。
- 必须按照作者或许可方指定的方式归因。
数据集创建者
- Amar Subramanya (asubram@google.com)
- Sameer Singh (sameer@cs.umass.edu)
- Fernando Pereira (pereira@google.com)
- Andrew McCallum (mccallum@cs.umass.edu)
- Dave Orr (dmorr@google.com)
注意事项
- 数据集自动从网络创建,可能包含一定程度的噪声。




