ArabicMTEB
收藏资源简介:
ArabicMTEB是一个全面的阿拉伯语文本嵌入基准数据集,由不列颠哥伦比亚大学开发。该数据集涵盖了8个不同的任务类别,包括跨语言检索、分类和语义相似性等,共包含94个数据集。数据集的内容丰富多样,包括标准阿拉伯语和多种方言,旨在评估文本嵌入模型在不同阿拉伯语环境中的表现。数据集的创建过程结合了人工生成和合成数据,确保了语言覆盖的全面性和多样性。ArabicMTEB的应用领域广泛,旨在解决阿拉伯语自然语言处理中的复杂问题,特别是在跨语言和跨文化环境中的文本理解和生成。
ArabicMTEB is a comprehensive Arabic text embedding benchmark dataset developed by the University of British Columbia. It covers 8 distinct task categories including cross-lingual retrieval, classification, semantic similarity and more, with a total of 94 datasets. The dataset has rich and diverse content, including Modern Standard Arabic and multiple dialects, aiming to evaluate the performance of text embedding models across different Arabic language scenarios. Its creation process combines manual curation and synthetic data generation, ensuring comprehensive and diverse linguistic coverage. ArabicMTEB has a wide range of application scenarios, and is designed to solve complex problems in Arabic natural language processing, especially text understanding and generation in cross-lingual and cross-cultural contexts.

- 1Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks不列颠哥伦比亚大学 · 2024年



