Chinese Multi-Document Question Answering Dataset (ChiMDQA)
收藏资源简介:
ChiMDQA数据集是一个涵盖学术、教育、金融、法律、医疗和新闻六个领域的长文本数据集,包含6068个经过严格筛选的高质量问答对,分为十个细粒度的类别。数据集旨在为中文文档问答任务提供高质量的、多样化的数据资源,适用于文档理解、知识提取和智能问答系统等多个NLP任务。数据集的构建过程包括数据收集、问答对生成、数据审查和验证统计等多个阶段。
The ChiMDQA Dataset is a long-text dataset covering six domains: academia, education, finance, law, healthcare, and journalism. It contains 6068 strictly curated high-quality question-answer pairs, which are divided into ten fine-grained categories. This dataset aims to provide high-quality and diversified data resources for Chinese document question answering tasks, and is applicable to multiple NLP tasks such as document understanding, knowledge extraction and intelligent question answering systems. The construction pipeline of the dataset includes multiple stages such as data collection, question-answer pair generation, data review and verification statistics.

- 1ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation北京交通大学, 北京邮电大学, 北京工业大学, 福建福昕软件有限公司 · 2025年



