cx-cmu/AgentWebBench-corpus
收藏资源简介:
AgentWebBench Corpus是一个为AgentWebBench基准测试预构建的密集检索语料库,用于评估多智能体在Agentic Web中的协调性能。该语料库基于ClueWeb22数据集的100个网站切片(约1840万份文档),包含文档嵌入向量(维度为1024)和FAISS索引(包括每个网站的独立索引和全局索引)。数据集不包含原始文本,仅提供向量和文档ID,需配合ClueWeb22原始数据使用。它适用于信息检索、密集检索等NLP任务,使用MIT许可证,语言为英文。
AgentWebBench Corpus is a pre-built dense-retrieval corpus for the AgentWebBench benchmark, designed for evaluating multi-agent coordination in Agentic Web over a realistic 100-website slice of ClueWeb22 (approximately 18.4 million documents). It includes embeddings (dimension 1024) and FAISS indices (per-website and global indices) but does not contain the original ClueWeb22 text. The corpus is used for tasks such as information retrieval and dense retrieval, under the MIT license, with English as the primary language.




