Urdu-Language Tourism Question-Answering Dataset for Gilgit-Baltistan
收藏资源简介:
This dataset contains an Urdu-language question-answering dataset focused on tourism in Gilgit-Baltistan, Pakistan. It comprises 900 question-answer pairs spanning three difficulty tiers: Easy (single-fact, single-source), Medium (multi-fact, single-source), and Hard (multi-source integration). Questions were sourced and cross-verified from visitgilgitbaltistan.gov.pk, travel blogs, and other publicly available sources, then translated and refined into Urdu by native speakers. A locked 150-question stratified subset (50 questions per difficulty tier) is included separately as the fixed evaluation set used in the accompanying study, which benchmarks small language models (SLMs) on retrieval-augmented generation (RAG) for Urdu tourism QA. Note: the underlying source/context corpus used for retrieval is not included in this release due to third-party copyright restrictions on the source material. Only the question-answer pairs are published here.



