遇见数据集

Question Benchmark Dataset for Valencia Tourism information RAG

收藏
Zenodo2025-10-18 更新2026-05-26 收录
官方服务:

资源简介:

Dataset Description This repository presents a domain-specific question-answering (QA) dataset designed for Retrieval-Augmented Generation (RAG) applications in the tourism domain, with a focus on Valencia, Spain. The dataset comprises 994 question-answer pairs grounded in a curated corpus of 84 Spanish-language documents pertaining to Valencia's tourism information. Corpus Composition The textual corpus has been systematically compiled from authoritative sources, including: Tourism-oriented blog articles Wikipedia entries related to Valencia Official tourism guides published by VisitValencia All source materials are provided in Spanish and have undergone standardized preprocessing procedures, including conversion to plain text format and coreference resolution to normalize entity mentions throughout the corpus. Dataset Structure The repository contains: documents.zip: Complete collection of 84 curated source documents QA dataset (dataset.csv): 994 question-answer pairs, each annotated with: Context window (relevant text excerpt containing the answer) Source document identifier Generation Methodology The question-answer pairs were generated programmatically using Google's Gemini 2.5 Flash model, ensuring each QA pair is contextually grounded in the source documentation. This approach guarantees factual consistency between generated questions, answers, and supporting textual evidence. Intended Applications This dataset is specifically designed for the development, training, and evaluation of domain-specific RAG systems targeting tourism information retrieval and question-answering tasks in Spanish-language contexts.

提供机构:
Zenodo
创建时间:
2025-10-18
二维码
社区交流群
二维码
科研交流群
商业服务