zelda
收藏资源简介:
ZELDA是一个用于实体消歧的基准数据集。由于ZELDA中没有开发集分割,我们将数据集的前90%用于训练,剩余的10%用于开发。对于实体字典,我们使用了由[Rücker和Akbik, 2025](https://github.com/flairNLP/VerbalizED)处理的维基百科页面ID和维基数据描述。 - **仓库:** [https://github.com/flairNLP/zelda](https://github.com/flairNLP/zelda) - **公开:** 是 - **来源:** [Kensho Derived Wikimedia Dataset](https://www.kaggle.com/datasets/kenshoresearch/kensho-derived-wikimedia-data) - **论文:** [ZELDA: A Comprehensive Benchmark for Supervised Entity Disambiguation](https://aclanthology.org/2023.eacl-main.151/) - **实体数量:** 821401
数据集概述:ZELDA
基本描述
ZELDA是一个用于实体消歧的基准数据集。该数据集是公开的。
来源与构成
- 数据来源:Kensho Derived Wikimedia Dataset。
- 实体数量:821,401个。
- 实体字典:使用经过Rücker and Akbik, 2025处理的维基百科页面ID和维基数据描述构建。
数据集划分与配置
数据集包含以下配置:
- 主数据集 (
dataset):- 训练集 (
train):data/train-00001-of-00001.jsonl - 验证集 (
validation):data/dev-00001-of-00001.jsonl - 注:由于原始ZELDA数据集不包含开发集,此处将前90%数据用于训练,剩余10%用于开发(验证)。
- 训练集 (
- 字典 (
dictionary):- 知识库 (
kb):dictionary/dictionary-00001-of-00001.jsonl
- 知识库 (
规模类别
数据集规模为:100K < n < 1M。
相关资源
- 代码仓库:https://github.com/flairNLP/zelda
- 论文:ZELDA: A Comprehensive Benchmark for Supervised Entity Disambiguation
引用信息
@inproceedings{milich2023zelda, title={{ZELDA}: A Comprehensive Benchmark for Supervised Entity Disambiguation}, author={Milich, Marcel and Akbik, Alan}, booktitle={{EACL} 2023, The 17th Conference of the European Chapter of the Association for Computational Linguistics}, year={2023} }




