MESED
收藏资源简介:
MESED是由清华大学创建的第一个大规模多模态实体集扩展数据集,包含14,489个来自维基百科的实体和434,675对图像-句子。该数据集设计了26个粗粒度和70个细粒度语义类别,用于评估模型在处理复杂实体如负实体、同义实体、多义实体和长尾实体时的表现。MESED旨在通过多模态信息提高实体表示的准确性,解决单一文本模态在实体扩展任务中的局限性,并应用于知识挖掘、网络搜索、分类体系构建和知识图谱等领域。
MESED is the first large-scale multimodal entity set expansion dataset developed by Tsinghua University, comprising 14,489 entities sourced from Wikipedia and 434,675 image-sentence pairs. The dataset features 26 coarse-grained and 70 fine-grained semantic categories, designed to evaluate model performance when handling complex entities including negative entities, synonymous entities, polysemous entities and long-tail entities. MESED aims to improve the accuracy of entity representation via multimodal information, address the limitations of single-text modality in entity expansion tasks, and support applications in domains such as knowledge mining, web search, classification system construction and knowledge graphs.




