mteb/MultilingualNanoQuoraRetrieval
收藏资源简介:
MultilingualNanoQuoraRetrieval是一个多语言文本检索数据集,属于大规模文本嵌入基准(MTEB)。该数据集是QuoraRetrieval数据集的一个较小子集,基于Quora平台上被标记为重复的问题构建。核心任务是在给定一个问题的情况下,从语料库中检索出其他语义相同或相似的重复问题。数据集覆盖11种语言:阿拉伯语、德语、英语、西班牙语、法语、意大利语、日语、韩语、挪威语、葡萄牙语和瑞典语。每种语言都包含三个部分:语料库(corpus,即文档集合)、查询(queries,即问题)和相关性判断(qrels,即查询与文档的匹配关系)。该数据集主要用于评估多语言文本嵌入模型在跨语言检索任务上的性能,特别适用于社交领域的问题去重和语义匹配场景。
MultilingualNanoQuoraRetrieval is a multilingual text retrieval dataset belonging to the Massive Text Embedding Benchmark (MTEB). It is a smaller subset of the QuoraRetrieval dataset, constructed based on questions marked as duplicates on the Quora platform. Its core task is to retrieve semantically identical or similar duplicate questions from the corpus given a query question. The dataset covers 11 languages: Arabic, German, English, Spanish, French, Italian, Japanese, Korean, Norwegian, Portuguese, and Swedish. Each language contains three components: corpus (document collection), queries (i.e., target questions), and qrels (query-document relevance matching relationships). This dataset is primarily used to evaluate the performance of multilingual text embedding models on cross-lingual retrieval tasks, and is particularly suitable for question deduplication and semantic matching scenarios in the social domain.




