Twitter2015-Urdu
收藏资源简介:
Twitter2015-Urdu数据集是首个为乌尔都语多模态命名实体识别(MNER)设计的MNER数据集,由Twitter2015英语数据集翻译和标注而成,确保了文本和图像在文化和语言上的相关性。该数据集经过精心设计,以支持严格的实验,并促进乌尔都语MNER研究的进展。该数据集的创建过程包括数据收集、数据预处理、翻译和审查、分词、数据标注以及质量控制与验证等关键步骤。该数据集的发布旨在解决低资源语言如乌尔都语在MNER领域的挑战,并为未来研究提供基准数据集。
The Twitter2015-Urdu dataset is the first multi-modal named entity recognition (MNER) dataset specifically tailored for the Urdu language. Developed through translation and annotation of the English Twitter2015 dataset, it guarantees cultural and linguistic relevance between the textual and visual components. This dataset is meticulously designed to support rigorous experimental research and promote the advancement of Urdu MNER studies. The construction of the dataset involves several key stages: data collection, data preprocessing, translation and review, tokenization, data annotation, as well as quality control and validation. The release of this dataset aims to address the challenges faced by low-resource languages such as Urdu in the MNER field, while providing a benchmark dataset for future research.

- 1A Benchmark Dataset and a Framework for Urdu Multimodal Named Entity Recognition北京化工大学信息科学与技术学院 · 2025年



