mteb/MemGovern
收藏资源简介:
该数据集是一个用于文本检索任务的英文单语数据集,派生自KaLM-Embedding/LMEB数据集。它包含多个开源项目的子数据集,每个项目作为一个独立配置,涵盖如Azure SDK、Flocker、DataDog、TypeScript、React、Kubernetes等知名项目。每个配置包括三个部分:语料库(corpus,包含文档ID、文本内容和标题)、查询(queries,包含查询ID和文本)以及相关性判断(qrels,包含查询ID、语料库ID和相关性分数)。所有数据仅提供测试分割,用于评估信息检索或文本检索系统的性能。数据集采用MIT许可证。
This dataset is an English monolingual dataset for text retrieval tasks, derived from the KaLM-Embedding/LMEB dataset. It consists of multiple sub-datasets from various open-source projects, each configured independently, covering well-known projects such as Azure SDK, Flocker, DataDog, TypeScript, React, Kubernetes, and more. Each configuration includes three components: a corpus (with document ID, text content, and title), queries (with query ID and text), and relevance judgments (qrels, with query ID, corpus ID, and relevance score). All data is provided only in a test split, designed for evaluating the performance of information retrieval or text retrieval systems. The dataset is licensed under MIT.



