imabari_wiki_qa_v4_validated
收藏资源简介:
Imabari QA v4 — Validated 是一个用于监督微调(SFT)的日语问答数据集。该数据集由两个独立构建且经过验证的变体合并而成:程序化验证数据集(4,146个样本)和人工验证数据集(4,078个样本)。源语料库为 ikedachin/imabari_wiki_cpt_v3,QA生成模型为 Qwen3.8-27B-NVFP4。合并后的数据集共包含8,224个样本,其中训练集7,401个样本,验证集823个样本(约9:1比例)。数据字段包括:qa_id(唯一UUID)、question(问题)、thinking(生成的推理过程)、answer(最终答案)、eval(QA评估值)、qa_generator(生成模型标识)、messages(ChatML格式的SFT对话字段)、source_files(源文件列表)、id(原始文档ID)和chunk_index(原始数据块索引)。数据集主要特点包括:日语问答、涵盖今治市和爱媛县等地域知识、合成推理数据、两种验证方法(程序化和人工)的整合、支持多种SFT格式(Chat SFT、推理SFT等)。预期用途包括日语LLM监督微调、指令微调、推理SFT、LoRA/QLoRA、区域特化LLM、区域知识问答、合成QA/推理研究、数据验证研究以及不同验证方法效果的对比实验。需要特别注意:本数据集的“验证”标记仅表示每个样本通过了至少一种验证方法(程序或人工),并非同时通过两者;数据为合成数据,可能包含幻觉、事实错误、推理错误、偏见和生成模型特定模式;程序化验证和人工验证标准不同,整体数据集验证标准不均匀;数据集不应用于医疗、法律、金融、防灾等高风险领域作为唯一信息源。数据集基于CC BY-SA 4.0许可发布。
Imabari QA v4 — Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT). It is created by merging two independently constructed and validated variants: a programmatically validated set (4,146 samples) and a human-validated set (4,078 samples). The source corpus is ikedachin/imabari_wiki_cpt_v3, and the QA generation model is Qwen3.8-27B-NVFP4. The merged dataset contains 8,224 samples, split into 7,401 training samples and 823 validation samples (approximately 9:1 ratio). Data fields include: qa_id (unique UUID), question, thinking (generated reasoning process), answer (final answer), eval (QA evaluation score), qa_generator (generator model identifier), messages (ChatML-formatted SFT conversation field), source_files (list of source files), id (original document ID), and chunk_index (original chunk index). Key features include: Japanese Q&A, local knowledge of Imabari City and Ehime Prefecture, synthetic reasoning data, integration of two validation methods (programmatic and human), and support for multiple SFT formats (Chat SFT, reasoning SFT, etc.). Intended uses include supervised fine-tuning of Japanese LLMs, instruction tuning, reasoning SFT, LoRA/QLoRA, region-specific LLMs, local knowledge Q&A, synthetic QA/reasoning research, data validation research, and comparative experiments on different validation methods. Important notes: The validated label only indicates that each sample passed at least one validation method (programmatic or human), not both; the data is synthetic and may contain hallucinations, factual errors, reasoning errors, biases, and generative model-specific patterns; programmatic and human validation standards differ, leading to uneven validation criteria; the dataset should not be used as a sole information source in high-risk domains such as medical, legal, financial, or disaster prevention. The dataset is released under the CC BY-SA 4.0 license.





