mteb/MultilingualNanoSCIDOCSRetrieval
收藏资源简介:
MultilingualNanoSCIDOCSRetrieval是一个多语言文本检索数据集,属于MTEB(大规模文本嵌入基准)的一部分。该数据集是NanoFiQA2018的较小子集,而NanoFiQA2018又源自SciDocs,后者是一个包含七个文档级任务的评估基准,涵盖引用预测、文档分类和推荐等任务。任务类别为文本到文本检索,领域涉及学术、书面和非虚构内容。数据集支持11种语言:阿拉伯语、德语、英语、法语、意大利语、日语、韩语、挪威语、葡萄牙语、西班牙语和瑞典语。数据来源于两个源数据集:zeta-alpha-ai/NanoSCIDOCS和LiquidAI/nanobeir-multilingual-extended,并包含测试集,具体包括语料库、查询和相关度评分(qrels)等组件,用于评估嵌入模型在多语言环境下的检索性能。
MultilingualNanoSCIDOCSRetrieval is a multilingual text retrieval dataset that forms part of the MTEB (Massive Text Embedding Benchmark). It is a smaller subset of NanoFiQA2018, which itself is derived from SciDocs—a benchmark evaluation suite comprising seven document-level tasks covering citation prediction, document classification, recommendation, and more. The task falls under the text-to-text retrieval category, with the domain spanning academic, written, and non-fiction content. The dataset supports 11 languages: Arabic, German, English, French, Italian, Japanese, Korean, Norwegian, Portuguese, Spanish, and Swedish. It is sourced from two original datasets: zeta-alpha-ai/NanoSCIDOCS and LiquidAI/nanobeir-multilingual-extended, and includes a test set with components such as corpus, queries, and relevance scores (qrels) to evaluate the retrieval performance of embedding models in multilingual settings.




