遇见数据集

trec-ragtime/ragtime2

收藏
Hugging Face2026-05-06 更新2026-05-31 收录
官方服务:

资源简介:

RAGTIME2数据集是为TREC RAGTIME Track 2026设计的多语言文档集合,用于检索增强生成(RAG)任务。该任务要求系统从所有四种语言(阿拉伯语、英语、西班牙语、中文)中检索相关文档,并生成带有引用的响应。数据集包含从Common Crawl News中提取的文档,采样时间范围为2021年8月1日至2024年7月31日,每天有均匀数量的文档。每种语言有1,000,095个文档,并提供了基于HLTCOE训练的Sockeye模型的机器翻译版本。文档按语言分为四个.jsonl文件,但旨在作为一个整体使用。

The RAGTIME2 dataset contains documents for the TREC RAGTIME Track 2026, a multilingual RAG task that expects systems to retrieve relevant documents from all four languages (Arabic, English, Spanish, Chinese) and synthesize a response with citation. The documents are extracted from Common Crawl News and sampled between August 1, 2021, and July 31, 2024, with an even number of documents every day. Each language has 1,000,095 documents, and machine translations using a Sockeye model trained by HLTCOE are also released. The documents are separated by language into four .jsonl files but are intended to be used as a whole set.

提供机构:
trec-ragtime
二维码
社区交流群
二维码
科研交流群
商业服务