Urhobo Fact-Checking Dataset with Evidence URLs
收藏资源简介:
Dataset title: Urhobo Fact-Checking Dataset with Evidence URLs Overview This dataset is a fact-checking and evidence-linked question-and-answer resource in Urhobo. Each record consists of a question and answer in Urhobo, together with an evidence URL intended to support the answer or the information presented in it.The dataset contains 178 records and 8 fields. All records use `urh_latn` in the `Language` field and `Nigeria` in the `Country/Region` field. English translations are available for both the questions and answers. Dataset structure The eight fields are:ID — unique identifier for each record.Language — language identifier for the source-language content; the observed value is `urh_latn`.Country/Region — country or region associated with the record; the observed value is `Nigeria`.Question — original question written in Urhobo.Answer — corresponding answer written in Urhobo.Evidence_Url — URL intended to provide evidence or supporting information for the answer.Question Translation — English translation of the Urhobo question, human-corrected during dataset preparation.Answer Translation — English translation of the Urhobo answer, human-corrected during dataset preparation. Content and thematic scope The questions and answers cover a broad range of subject areas. Topics represented in the dataset include Urhobo language and culture, naming practices, folklore, festivals, traditional practices, religion, food, music and social life; Ijaw and Rivers State cultural practices; education and language revitalization; Nigerian government, public policy, health, agriculture, security and current affairs; sports; artificial intelligence and work; personal finance; and gender and social issues.The collection therefore combines cultural and heritage-oriented information with general-interest and contemporary factual questions.Evidence sourcesEvery record contains an `Evidence_Url`. Across the 178 records, there are 125 unique URLs. The linked sources are diverse and may include academic publications, government and public-health resources, news organizations, reference materials, educational resources, social-media pages, and video platforms.The evidence URL identifies the source associated with a record. Researchers may evaluate the authority, accessibility, persistence, and suitability of individual sources according to the requirements of their study.Translation fieldsThe `Question Translation` and `Answer Translation` fields provide English translations of the original Urhobo content. These translations are the human-corrected translations included with the dataset and are intended to facilitate multilingual analysis, evaluation, and use by researchers who do not read Urhobo.The `Question Translation` field is populated for all 178 records. The `Answer Translation` field is populated for 177 records; one record has no answer translation. Data characteristics The dataset preserves the original Urhobo question-and-answer content alongside the English translations and associated evidence URLs. Record identifiers are sequential from 1 to 178. Some questions may be repeated or closely related, reflecting the contents of the collection. Intended uses This dataset may support research in: multilingual and low-resource natural language processing; fact-grounded and evidence-linked question answering; retrieval and evidence-aware language models; machine translation and translation post-editing evaluation; Urhobo language technology; African language data curation and benchmarking; cultural and linguistic heritage research; and development of language resources for underrepresented languages. The dataset should be interpreted as a collection of questions and answers paired with supporting evidence URLs and English translations. The presence of an evidence URL indicates the source associated with a record; it should not be taken as a guarantee that the linked source remains available or that it independently establishes every aspect of a claim.Users should preserve the original source-language content when conducting multilingual or linguistic analyses and should evaluate evidence sources according to the standards appropriate to their research.



