HITL-RAG FAITHFULNESS EVALUATION DATASET FOR CULTURAL TOURISM INFORMATION
收藏资源简介:
This dataset supports accuracy evaluation in Retrieval-Augmented Generation (RAG) systems applied to cultural tourism information, with an explicit focus on Human-in-the-Loop (HITL) governance. The dataset is designed as an evaluation artifact, not a performance benchmark, emphasizing traceability, foundational accuracy, and decision transparency. The dataset consists of structured query-response evaluations, including retrieved text snippets, candidate answers generated by a RAG-based language model, reference annotations generated by an LLM-based rater, automated similarity metrics, and final evaluation decisions by humans. Accuracy is assessed based on curated cultural knowledge snippets, while system-level backup and failure handling mechanisms are explicitly separated from content-level foundations. A structured HITL error taxonomy (E0–E10) is implemented to distinguish accurate responses, controlled rejections, content errors, and system failures. Repeated queries are intentionally included to observe consistency across the evaluation cycle, enabling generative variation analysis while maintaining stable classification results under human judgment. Version management follows a snapshot-based policy to ensure reproducibility and comparability between versions. This release (Version 1.2) contains no substantial changes except for image quality improvements and the addition of a metadata table. The release of Version 1.2 remains confidential and will be made publicly available upon publication in the relevant journal.



