nepali-qa-9k
收藏资源简介:
Nepali-Qa-9k是一个合成生成的尼泊尔语问答对数据集,专门用于训练和评估检索模型、嵌入模型、重排序器以及FAQ(常见问题解答)风格的搜索系统。数据集包含总计9,000条记录,分为7,200条训练样本和1,800条测试样本。每个数据样本由四个字段构成:domain(领域,字符串类型),表示问题所属的类别;scenario(场景,字符串类型),描述具体情境;question(问题,字符串类型),即用自然尼泊尔语书写的用户查询;以及answer(答案,字符串类型),即对应的简洁、有帮助的回答。问题模拟了用户在FAQ页面、帮助中心、客服聊天机器人、搜索界面或信息检索系统中可能提出的真实世界查询。数据集涵盖了广泛的日常生活领域,包括健康、教育、银行、保险、政府服务、旅行、就业、技术、购物、法律程序、家庭问题、交通、移动应用、育儿和日常支持等。数据通过结构化提示过程生成,使用Gemma-4-E2B-IT模型,并基于预定义的领域和场景进行约束,以提高多样性和减少重复。该数据集适用于构建尼泊尔语检索模型、微调句子嵌入、开发FAQ搜索系统、训练问答匹配模型以及搭建客户支持或帮助中心检索管道等任务。需要注意的是,这是一个合成数据集,可能包含不完美、重复模式或对某些实际情况过于笼统的答案,且其答案不应被视为专业的医疗、法律、金融或政府建议。
Nepali-Qa-9k is a synthetically generated Nepali question-answer pair dataset, specifically designed for training and evaluating retrieval models, embedding models, re-rankers, and FAQ-style search systems. The dataset contains a total of 9,000 records, divided into 7,200 training samples and 1,800 test samples. Each data sample consists of four fields: domain (string type), indicating the category of the question; scenario (string type), describing the specific context; question (string type), which is the user query written in natural Nepali; and answer (string type), the corresponding concise and helpful response. The questions simulate real-world queries that users might pose on FAQ pages, help centers, customer service chatbots, search interfaces, or information retrieval systems. The dataset covers a wide range of daily life domains, including health, education, banking, insurance, government services, travel, employment, technology, shopping, legal procedures, family issues, transportation, mobile applications, parenting, and daily support. The data is generated through a structured prompting process using the `Gemma-4-E2B-IT` model, constrained by predefined domains and scenarios to enhance diversity and reduce repetition. This dataset is suitable for tasks such as building Nepali retrieval models, fine-tuning sentence embeddings, developing FAQ search systems, training question-answer matching models, and constructing customer support or help center retrieval pipelines. It is important to note that this is a synthetic dataset and may contain imperfections, repetitive patterns, or overly general answers for certain real-world situations, and its answers should not be considered as professional medical, legal, financial, or governmental advice.
数据集名称
Nepali-Qa-9k
数据集概述
这是一个合成的尼泊尔语问答数据集,包含9,000个样本,专为训练和评估检索模型、嵌入模型、重排序模型以及FAQ风格搜索系统而设计。
数据集规模
- 样本总数: 9,000对问答对
- 数据集大小: 6.47 MB (下载大小: 1.69 MB)
数据拆分
- 训练集 (train): 7,200个样本 (5.18 MB)
- 测试集 (test): 1,800个样本 (1.29 MB)
数据特征
每个样本包含4个字段:
- domain: 领域/类别标签 (字符串)
- scenario: 场景描述 (字符串)
- question: 用户问题 (字符串)
- answer: 答案 (字符串)
语言
- 尼泊尔语
任务类别
- 问答 (question-answering)
数据来源与生成方式
- 使用
Gemma-4-E2B-IT模型通过结构化提示生成 - 每个生成样本基于预定义的领域和场景,以提高多样性和减少重复
- 问题以自然尼泊尔语编写,模拟FAQ页面、帮助中心、客服聊天机器人等真实场景
覆盖领域
数据集涵盖多个日常领域,包括:
- स्वास्थ्य (健康)
- शिक्षा (教育)
- बैंकिङ (银行)
- बीमा (保险)
- सरकारी सेवा (政府服务)
- यात्रा (旅行)
- रोजगारी (就业)
- अभिभावकत्व (育儿)
- प्रविधि (技术)
- किनमेल (购物)
- कानुनी प्रक्रिया (法律程序)
- घरायसी समस्या (家庭问题)
- यातायात (交通)
- मोबाइल एप (移动应用)
- दैनिक जीवन (日常生活)
预期用途
- 训练尼泊尔语检索模型
- 微调句子嵌入模型
- 构建FAQ搜索系统
- 训练问答匹配模型
- 开发客服或帮助中心检索流程
局限性
- 这是一个合成数据集,可能存在不完美之处、重复模式或过于笼统的回答
- 不应用于专业医疗、法律、金融、政府等敏感领域的决策,需核实重要信息
引用信息
bibtex @dataset{subedi_nepali_nli_dataset, author = {Subedi, Sanjaya}, title = {Nepali Question Answer Dataset}, year = {2026}, publisher = {Hugging Face}, url = {https://huggingface.co/datasets/jangedoo/nepali-qa-9k} }





