ft-llm-2026-domain-specific-qa
收藏资源简介:
该数据集是一个多模态视觉问答数据集,包含32,484个训练样本。每个数据样本由以下要素构成:图像数据、自然语言问题、文本答案、问题类型标注、难度等级、推理过程说明、答案解释、来源PDF文档路径、页码位置、QA索引号、相关度评分、模型标识和子文件夹分类。数据集特别包含学术文档来源追踪(通过source_pdf和page_num字段)和答案质量评估指标(relevance_score),适用于视觉问答系统开发、多模态推理研究以及学术文档信息提取等任务。数据总量约为12.3GB,采用PDL 1.0许可证。
This dataset is a multimodal visual question answering (VQA) dataset containing 32,484 training samples. Each data sample comprises the following components: image data, natural language questions, textual answers, question type annotations, difficulty levels, reasoning process descriptions, answer explanations, source PDF document paths, page number positions, QA index numbers, relevance scores, model identifiers, and subfolder classifications. The dataset specifically provides academic document source tracking via the source_pdf and page_num fields, as well as answer quality evaluation metrics represented by the relevance_score field. It is suitable for tasks including visual question answering system development, multimodal reasoning research, and academic document information extraction. The total data volume is approximately 12.3 GB, and the dataset is released under the PDL 1.0 license.



