MEE
收藏资源简介:
MEE数据集是由俄勒冈大学计算机科学系创建的一个新型多语言事件抽取数据集,旨在解决非英语语言在事件抽取研究中的不足。该数据集包含超过50,000个事件提及,涵盖8种不同语系的语言,包括英语、西班牙语、葡萄牙语、波兰语、土耳其语、印地语、韩语和日语。MEE数据集全面标注了实体提及、事件触发词和事件参数,支持跨语言迁移学习评估。该数据集的应用领域包括问答系统、知识库扩充和文本摘要,旨在提高模型在多语言环境下的性能和泛化能力。
The MEE dataset is a novel multilingual event extraction dataset developed by the Department of Computer Science at the University of Oregon, aimed at addressing the underrepresentation of non-English languages in event extraction research. This dataset contains over 50,000 event mentions, spanning 8 distinct language families including English, Spanish, Portuguese, Polish, Turkish, Hindi, Korean, and Japanese. The MEE dataset comprehensively annotates entity mentions, event triggers, and event arguments, supporting cross-lingual transfer learning evaluation. Its application areas include question answering systems, knowledge base augmentation, and text summarization, with the objective of enhancing model performance and generalization ability in multilingual settings.




