ArabicMTEB
收藏资源简介:
ArabicMTEB是一个全面的阿拉伯语文本嵌入基准,旨在评估跨语言、多方言、多领域和多文化的阿拉伯语文本嵌入性能。该数据集包含94个数据集,涵盖8个不同的任务,包括检索、分类和语义相似性等。数据集的内容丰富多样,包括标准阿拉伯语和各种方言,以及不同领域的文本。创建过程涉及使用人工生成和合成数据,确保了数据集的广泛覆盖和多样性。该数据集主要应用于阿拉伯语自然语言处理领域,旨在解决阿拉伯语特有的语言和文化复杂性问题。
ArabicMTEB is a comprehensive Arabic text embedding benchmark developed to evaluate the performance of Arabic text embeddings across cross-lingual, multi-dialectal, multi-domain, and multi-cultural dimensions. This benchmark comprises 94 datasets covering 8 distinct task types, such as retrieval, classification, semantic similarity, and others. The datasets feature rich and diverse content, encompassing Modern Standard Arabic (MSA), various Arabic dialects, and texts from diverse domains. Its development process incorporates both human-generated and synthetic data, ensuring extensive coverage and adequate diversity of the benchmark datasets. Primarily utilized in the field of Arabic natural language processing (NLP), this benchmark aims to address the unique linguistic and cultural complexities inherent to the Arabic language.




