遇见数据集

adithya7/xlel_wd_dictionary

收藏
Hugging Face2022-07-01 更新2024-03-04 收录
官方服务:

资源简介:

--- annotations_creators: - found language_creators: - found language: - af - ar - be - bg - bn - ca - cs - da - de - el - en - es - fa - fi - fr - he - hi - hu - id - it - ja - ko - ml - mr - ms - nl - 'no' - pl - pt - ro - ru - si - sk - sl - sr - sv - sw - ta - te - th - tr - uk - vi - zh license: - cc-by-4.0 multilinguality: - multilingual pretty_name: XLEL-WD is a multilingual event linking dataset. This supplementary dataset contains a dictionary of event items from Wikidata. The descriptions for Wikidata event items are taken from the corresponding multilingual Wikipedia articles. size_categories: - 10K<n<100K source_datasets: - original task_categories: [] task_ids: [] --- # Dataset Card for XLEL-WD-Dictionary ## Table of Contents - [Table of Contents](#table-of-contents) - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards) - [Languages](#languages) - [Dataset Structure](#dataset-structure) - [Data Instances](#data-instances) - [Data Fields](#data-fields) - [Data Splits](#data-splits) - [Dataset Creation](#dataset-creation) - [Curation Rationale](#curation-rationale) - [Source Data](#source-data) - [Annotations](#annotations) - [Personal and Sensitive Information](#personal-and-sensitive-information) - [Considerations for Using the Data](#considerations-for-using-the-data) - [Social Impact of Dataset](#social-impact-of-dataset) - [Discussion of Biases](#discussion-of-biases) - [Other Known Limitations](#other-known-limitations) - [Additional Information](#additional-information) - [Dataset Curators](#dataset-curators) - [Licensing Information](#licensing-information) - [Citation Information](#citation-information) - [Contributions](#contributions) ## Dataset Description - **Homepage:** <https://github.com/adithya7/xlel-wd> - **Repository:** <https://github.com/adithya7/xlel-wd> - **Paper:** <https://arxiv.org/abs/2204.06535> - **Leaderboard:** N/A - **Point of Contact:** Adithya Pratapa ### Dataset Summary XLEL-WD is a multilingual event linking dataset. This supplementary dataset contains a dictionary of event items from Wikidata. The descriptions for Wikidata event items are taken from the corresponding multilingual Wikipedia articles. ### Supported Tasks and Leaderboards This dictionary can be used as a part of the event linking task. ### Languages This dataset contains text from 44 languages. The language names and their ISO 639-1 codes are listed below. For details on the dataset distribution for each language, refer to the original paper. | Language | Code | Language | Code | Language | Code | Language | Code | | -------- | ---- | -------- | ---- | -------- | ---- | -------- | ---- | | Afrikaans | af | Arabic | ar | Belarusian | be | Bulgarian | bg | | Bengali | bn | Catalan | ca | Czech | cs | Danish | da | | German | de | Greek | el | English | en | Spanish | es | | Persian | fa | Finnish | fi | French | fr | Hebrew | he | | Hindi | hi | Hungarian | hu | Indonesian | id | Italian | it | | Japanese | ja | Korean | ko | Malayalam | ml | Marathi | mr | | Malay | ms | Dutch | nl | Norwegian | no | Polish | pl | | Portuguese | pt | Romanian | ro | Russian | ru | Sinhala | si | | Slovak | sk | Slovene | sl | Serbian | sr | Swedish | sv | | Swahili | sw | Tamil | ta | Telugu | te | Thai | th | | Turkish | tr | Ukrainian | uk | Vietnamese | vi | Chinese | zh | ## Dataset Structure ### Data Instances Each instance in the `label_dict.jsonl` file follows the below template, ```json { "label_id": "830917", "label_title": "2010 European Aquatics Championships", "label_desc": "The 2010 European Aquatics Championships were held from 4–15 August 2010 in Budapest and Balatonfüred, Hungary. It was the fourth time that the city of Budapest hosts this event after 1926, 1958 and 2006. Events in swimming, diving, synchronised swimming (synchro) and open water swimming were scheduled.", "label_lang": "en" } ``` ### Data Fields | Field | Meaning | | ----- | ------- | | `label_id` | Wikidata ID | | `label_title` | Title for the event, as collected from the corresponding Wikipedia article | | `label_desc` | Description for the event, as collected from the corresponding Wikipedia article | | `label_lang` | language used for the title and description | ### Data Splits This dictionary has a single split, `dictionary`. It contains 10947 event items from Wikidata and a total of 114834 text descriptions collected from multilingual Wikipedia articles. ## Dataset Creation ### Curation Rationale This datasets helps address the task of event linking. KB linking is extensively studied for entities, but its unclear if the same methodologies can be extended for linking mentions to events from KB. Event items are collected from Wikidata. ### Source Data #### Initial Data Collection and Normalization A Wikidata item is considered a potential event if it has spatial and temporal properties. The final event set is collected after post-processing for quality control. #### Who are the source language producers? The titles and descriptions for the events are written by Wikipedia contributors. ### Annotations #### Annotation process This dataset was automatically compiled from Wikidata. It was post-processed to improve data quality. #### Who are the annotators? Wikidata and Wikipedia contributors. ### Personal and Sensitive Information [More Information Needed] ## Considerations for Using the Data ### Social Impact of Dataset [More Information Needed] ### Discussion of Biases [More Information Needed] ### Other Known Limitations This dictionary primarily contains eventive nouns from Wikidata. It does not include other event items from Wikidata such as disease outbreak (Q3241045), military offensive (Q2001676), war (Q198), etc., ## Additional Information ### Dataset Curators The dataset was curated by Adithya Pratapa, Rishubh Gupta and Teruko Mitamura. The code for collecting the dataset is available at [Github:xlel-wd](https://github.com/adithya7/xlel-wd). ### Licensing Information XLEL-WD dataset is released under [CC-BY-4.0 license](https://creativecommons.org/licenses/by/4.0/). ### Citation Information ```bib @article{pratapa-etal-2022-multilingual, title = {Multilingual Event Linking to Wikidata}, author = {Pratapa, Adithya and Gupta, Rishubh and Mitamura, Teruko}, publisher = {arXiv}, year = {2022}, url = {https://arxiv.org/abs/2204.06535}, } ``` ### Contributions Thanks to [@adithya7](https://github.com/adithya7) for adding this dataset.

提供机构:
adithya7
原始信息汇总

数据集概述

数据集名称

  • 名称: XLEL-WD
  • 别名: XLEL-WD-Dictionary

数据集描述

数据集摘要

  • 描述: XLEL-WD是一个多语言事件链接数据集,包含从Wikidata收集的事件项字典。事件项的描述来自对应的多语言维基百科文章。

支持的任务

  • 任务: 事件链接

语言

  • 数量: 44种语言
  • 列表: 包括Afrikaans, Arabic, Belarusian, Bulgarian, Bengali, Catalan, Czech, Danish, German, Greek, English, Spanish, Persian, Finnish, French, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Malayalam, Marathi, Malay, Dutch, Norwegian, Polish, Portuguese, Romanian, Russian, Sinhala, Slovak, Slovene, Serbian, Swedish, Swahili, Tamil, Telugu, Thai, Turkish, Ukrainian, Vietnamese, Chinese等。

数据集结构

数据实例

  • 文件: label_dict.jsonl
  • 结构: 每个实例包含label_id, label_title, label_desc, label_lang四个字段。

数据字段

  • label_id: Wikidata ID
  • label_title: 事件标题,来自对应的维基百科文章
  • label_desc: 事件描述,来自对应的维基百科文章
  • label_lang: 标题和描述使用的语言

数据分割

  • 分割: dictionary
  • 数量: 包含10947个Wikidata事件项和114834个文本描述。

数据集创建

数据来源

  • 来源: Wikidata
  • 语言生产者: 维基百科贡献者

注释过程

  • 过程: 自动从Wikidata收集并进行后处理以提高数据质量。
  • 注释者: 维基百科和Wikidata贡献者

许可证

  • 类型: CC-BY-4.0

贡献者

  • 主要贡献者: Adithya Pratapa, Rishubh Gupta, Teruko Mitamura
搜集汇总
数据集介绍
adithya7/xlel_wd_dictionary 数据集图片
构建方式
在知识图谱与自然语言处理交叉领域中,事件链接任务致力于将文本中提及的事件与知识库中的结构化事件条目进行关联。XLEL-WD-Dictionary数据集正是为此任务量身打造的多语言事件词典资源。该数据集以维基数据(Wikidata)为基石,通过筛选具备空间与时间属性的事件条目,并经过严格的质量控制后处理,最终汇集了10,947个事件项。每个事件的标题与描述信息均从对应语言版本的维基百科文章中自动提取,覆盖了44种语言,共包含114,834条文本描述,形成了一部横跨多语种的综合性事件词典。
使用方法
在应用层面,该数据集以JSONL格式存储,每个实例包含四个字段:label_id(维基数据ID)、label_title(事件标题)、label_desc(事件描述)及label_lang(语言代码)。研究人员可直接将其作为事件链接任务中的知识库词典使用,通过匹配文本中提及的事件与词典中的条目实现链接。数据集仅包含一个名为“dictionary”的单一划分,便于直接加载与集成。使用时需注意,该词典主要涵盖事件性名词,不包含所有维基数据中的事件类型,因此在实际应用中需结合具体任务需求进行筛选或补充。
背景与挑战
背景概述
事件链接(Event Linking)作为知识库链接任务的重要分支,旨在将自然语言文本中提及的事件与结构化知识库中的事件条目进行精准关联。然而,相较于实体链接的成熟研究,事件链接因事件本身的复杂时空属性与语义多样性而面临独特挑战。在此背景下,由Adithya Pratapa、Rishubh Gupta和Teruko Mitamura于2022年创建的多语言事件链接数据集XLEL-WD应运而生,其配套词典数据集adithya7/xlel_wd_dictionary应运而生。该词典从Wikidata中提取了10,947个事件条目,并利用多语言维基百科文章为其提供了114,834条跨语言描述文本,覆盖44种语言,为多语言事件链接研究提供了关键的基础资源。该数据集以CC-BY-4.0许可发布,其核心研究问题在于探索如何将实体链接的成功方法论迁移至事件领域,从而推动知识库链接范式的拓展。
当前挑战
该数据集所面临的挑战主要源于事件链接任务的固有难点与数据构建过程的复杂性。在领域问题层面,事件本身具有动态演变性、时空依赖性和类别多样性(如战争、疫情、体育赛事等),导致事件与实体在知识库中的表示模式差异显著,现有实体链接模型难以直接迁移。此外,多语言场景下事件描述的跨语言语义对齐与歧义消解进一步加剧了任务难度。在构建过程中,数据集需从Wikidata海量条目中自动筛选具备时空属性的潜在事件,并通过后处理质量控制剔除噪声,但受限于自动标注的准确性,词典仍主要收录事件性名词,而遗漏了如疾病爆发(Q3241045)、军事行动(Q2001676)等非典型事件类别,导致事件覆盖范围存在偏差。同时,多语言描述的质量参差不齐,部分语言条目稀疏,可能引入语言间的不平衡性,影响下游模型的泛化能力。
常用场景
经典使用场景
XLEL-WD-Dictionary作为多语言事件链接任务的核心资源,为从非结构化文本中识别并链接到Wikidata中结构化事件知识提供了关键支撑。该数据集通过整合44种语言的维基百科描述,构建了涵盖10947个事件条目的字典,每个条目包含事件标题、描述及语言标签。研究者可将其作为实体链接范式的扩展,用于评估跨语言事件提及与知识库中事件项的映射能力,尤其适用于需要处理时空属性事件的复杂场景。
解决学术问题
该数据集有效突破了传统知识库链接研究聚焦于实体而忽视事件链接的局限。它解决了如何从海量多语言文本中自动识别事件提及并关联至Wikidata事件项的学术难题,为事件抽取、跨语言信息检索和知识图谱补全提供了标准化基准。通过提供事件性名词的字典化表示,它推动了事件链接从单语言向多语言的跨越,显著提升了跨语言事件知识对齐的精确性与鲁棒性。
实际应用
在实际应用中,该数据集可赋能多语言新闻聚合系统,自动将来自不同语种的新闻报道链接至统一的事件知识库,从而支持跨语言事件追踪与趋势分析。此外,它可用于智能问答系统,帮助用户通过自然语言查询快速定位特定历史事件(如体育赛事、自然灾害)的详细信息,或辅助数字人文研究中对全球性事件的多语种文献挖掘与关联分析。
数据集最近研究
最新研究方向
在知识图谱与自然语言处理交叉领域,事件链接任务正逐渐从实体链接的成熟范式向更具动态性与复杂性的方向演进。XLEL-WD数据集作为首个大规模多语言事件链接资源,通过整合Wikidata中具有时空属性的事件条目,并辅以多语种维基百科描述,为跨语言事件知识对齐提供了关键支撑。当前前沿研究聚焦于利用该数据集构建事件驱动的神经链接模型,探索事件提及与知识库中结构化事件项之间的语义映射机制。随着全球信息网络对实时事件理解需求的激增,该数据集在跨语言新闻分析、历史事件关联挖掘及多模态事件推理等热点场景中展现出重要价值,其精心筛选的44种语言覆盖与质量控制流程,为打破少数主流语言主导的事件知识壁垒奠定了坚实基础,推动了事件级语义理解从单一语言向普适性多语言范式的跨越。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务