HITL-RAG Faithfulness Evaluation Dataset for Cultural Tourism Information (Version 1.1)
收藏资源简介:
This dataset supports the evaluation of faithfulness in Retrieval-Augmented Generation (RAG) systems applied to cultural tourism information, with an explicit focus on Human-in-the-Loop (HITL) governance. It is designed as an evaluation artifact rather than a performance benchmark, emphasizing traceability, grounding fidelity, and decision transparency. The dataset consists of structured query–response evaluations, including retrieved text chunks, candidate answers generated by a RAG-based language model, reference explanations produced by an LLM-based judge, automatic similarity metrics, and final human evaluation decisions. Faithfulness is assessed with respect to curated cultural knowledge chunks, while system-level fallback and failure-trapping mechanisms are explicitly separated from content-level grounding. A structured HITL error taxonomy (E0–E10) is applied to distinguish faithful responses, controlled abstentions, content errors, and system failures. Repeated queries are intentionally included to observe consistency across evaluation cycles, allowing analysis of generative variation while maintaining stable classification outcomes under human judgment. Versioning follows a snapshot-based policy to ensure reproducibility and cross-version comparison. This release (Version 1.1) is deposited under embargo and will be made publicly accessible following completion of the associated journal publication.



