遇见数据集

MMNER

收藏
arXiv2025-09-30 收录
数据链接:
官方服务:

资源简介:

该数据集是一个大规模的多语言和多模态命名实体识别(NER)数据集,包含四种语言(英语、法语、德语和西班牙语)的图像-文本对。该数据集涵盖四个类别(人物、地点、组织和杂项),共89,019个实体,并且已经通过高可靠性的手动标注(Cohen's kappa系数为0.96)。其规模达到33,965个图像-文本对,适用于多语言和多模态命名实体识别(Mmner)任务。

This dataset is a large-scale multilingual and multimodal named entity recognition (NER) dataset containing image-text pairs across four languages: English, French, German and Spanish. It covers four categories (person, location, organization and miscellaneous) with a total of 89,019 entities, and has been manually annotated with high reliability (Cohen's kappa coefficient of 0.96). Comprising 33,965 image-text pairs, this dataset is suitable for multilingual and multimodal named entity recognition (Mmner) tasks.

二维码
社区交流群
二维码
科研交流群
商业服务