Evaluation Vector Database (ChromaDB) for Multi-Hop QA Generation based on The Guardian News Data
收藏资源简介:
This dataset contains a pre-computed ChromaDB vector database created for the evaluation of Retrieval-Augmented Generation (RAG) systems in multi-hop, multi-turn scenarios. It was developed as part of a Master's thesis extending the RAG-DIVE evaluation framework of Brehme et al. (2026) Data Source & Processing: Source: The Guardian Open Platform API. Timeframe: 01 January 2026 to 20 April 2026 (collected to ensure a strict zero-shot evaluation environment for LLMs with a 2025 parametric knowledge cutoff). Processing: The raw articles were processed using recursive character chunking (chunk size: 1000, overlap: 150) and embedded using Google's gemini-embedding-001 model. Chunks are explicitly enriched with metadata (headline, date, and a parent article_id) to technically support the stochastic multi-hop context progression mechanism. Copyright Notice & Access Restriction: The original text content and copyright of the news articles belong entirely to The Guardian. The data was accessed and processed exclusively for non-commercial, academic research purposes. Due to The Guardian's Terms of Service regarding the redistribution and sub-licensing of their content, this dataset cannot be published as Open Access. Access is provided strictly to thesis reviewers and authorised researchers for the purpose of scientific reproducibility. Please use the "Request Access" button and state your academic affiliation and purpose.



