遇见数据集

WHUIR/matinf

收藏
Hugging Face2024-01-18 更新2024-06-15 收录
官方服务:

资源简介:

--- paperswithcode_id: matinf pretty_name: Maternal and Infant Dataset dataset_info: - config_name: age_classification features: - name: question dtype: string - name: description dtype: string - name: label dtype: class_label: names: '0': 0-1岁 '1': 1-2岁 '2': 2-3岁 - name: id dtype: int32 splits: - name: train num_bytes: 33901977 num_examples: 134852 - name: test num_bytes: 9616194 num_examples: 38318 - name: validation num_bytes: 4869685 num_examples: 19323 download_size: 0 dataset_size: 48387856 - config_name: topic_classification features: - name: question dtype: string - name: description dtype: string - name: label dtype: class_label: names: '0': 产褥期保健 '1': 儿童过敏 '2': 动作发育 '3': 婴幼保健 '4': 婴幼心理 '5': 婴幼早教 '6': 婴幼期喂养 '7': 婴幼营养 '8': 孕期保健 '9': 家庭教育 '10': 幼儿园 '11': 未准父母 '12': 流产和不孕 '13': 疫苗接种 '14': 皮肤护理 '15': 宝宝上火 '16': 腹泻 '17': 婴幼常见病 - name: id dtype: int32 splits: - name: train num_bytes: 153326538 num_examples: 613036 - name: test num_bytes: 43877443 num_examples: 175363 - name: validation num_bytes: 21834951 num_examples: 87519 download_size: 0 dataset_size: 219038932 - config_name: summarization features: - name: description dtype: string - name: question dtype: string - name: id dtype: int32 splits: - name: train num_bytes: 181245403 num_examples: 747888 - name: test num_bytes: 51784189 num_examples: 213681 - name: validation num_bytes: 25849900 num_examples: 106842 download_size: 0 dataset_size: 258879492 - config_name: qa features: - name: question dtype: string - name: answer dtype: string - name: id dtype: int32 splits: - name: train num_bytes: 188047511 num_examples: 747888 - name: test num_bytes: 53708532 num_examples: 213681 - name: validation num_bytes: 26931809 num_examples: 106842 download_size: 0 dataset_size: 268687852 --- # Dataset Card for "matinf" ## Table of Contents - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards) - [Languages](#languages) - [Dataset Structure](#dataset-structure) - [Data Instances](#data-instances) - [Data Fields](#data-fields) - [Data Splits](#data-splits) - [Dataset Creation](#dataset-creation) - [Curation Rationale](#curation-rationale) - [Source Data](#source-data) - [Annotations](#annotations) - [Personal and Sensitive Information](#personal-and-sensitive-information) - [Considerations for Using the Data](#considerations-for-using-the-data) - [Social Impact of Dataset](#social-impact-of-dataset) - [Discussion of Biases](#discussion-of-biases) - [Other Known Limitations](#other-known-limitations) - [Additional Information](#additional-information) - [Dataset Curators](#dataset-curators) - [Licensing Information](#licensing-information) - [Citation Information](#citation-information) - [Contributions](#contributions) ## Dataset Description - **Homepage:** [https://github.com/WHUIR/MATINF](https://github.com/WHUIR/MATINF) - **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) - **Paper:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) - **Point of Contact:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) - **Size of downloaded dataset files:** 0.00 MB - **Size of the generated dataset:** 795.00 MB - **Total amount of disk used:** 795.00 MB ### Dataset Summary MATINF is the first jointly labeled large-scale dataset for classification, question answering and summarization. MATINF contains 1.07 million question-answer pairs with human-labeled categories and user-generated question descriptions. Based on such rich information, MATINF is applicable for three major NLP tasks, including classification, question answering, and summarization. We benchmark existing methods and a novel multi-task baseline over MATINF to inspire further research. Our comprehensive comparison and experiments over MATINF and other datasets demonstrate the merits held by MATINF. ### Supported Tasks and Leaderboards [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ### Languages [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ## Dataset Structure ### Data Instances #### age_classification - **Size of downloaded dataset files:** 0.00 MB - **Size of the generated dataset:** 48.39 MB - **Total amount of disk used:** 48.39 MB An example of 'validation' looks as follows. ``` This example was too long and was cropped: { "description": "\"6个月的时候去儿宝检查,医生说宝宝的分胯动作做的不好,说最好去儿童医院看看,但我家宝宝很好,感觉没有什么不正常啊,请教一下,分胯做的不好,有什么不好吗?\"...", "id": 88016, "label": 0, "question": "医生说宝宝的分胯动作不好" } ``` #### qa - **Size of downloaded dataset files:** 0.00 MB - **Size of the generated dataset:** 268.69 MB - **Total amount of disk used:** 268.69 MB An example of 'train' looks as follows. ``` This example was too long and was cropped: { "answer": "\"我一个同学的孩子就是发现了肾积水,治疗了一段时间,结果还是越来越多,没办法就打掉了。虽然舍不得,但是还是要忍痛割爱,不然以后孩子真的有问题,大人和孩子都受罪。不过,这个最后的决定还要你自己做,毕竟是你的宝宝。,、、、、\"...", "id": 536714, "question": "孕5个月检查右侧肾积水孩子能要吗?" } ``` #### summarization - **Size of downloaded dataset files:** 0.00 MB - **Size of the generated dataset:** 258.88 MB - **Total amount of disk used:** 258.88 MB An example of 'train' looks as follows. ``` This example was too long and was cropped: { "description": "\"宝宝有中度HIE,但原因未查明,这是他出生后脸上红的几道,嘴唇深红近紫,请问这是像缺氧的表现吗?\"...", "id": 173649, "question": "宝宝脸上红的几道嘴唇深红近紫是像缺氧的表现吗?" } ``` #### topic_classification - **Size of downloaded dataset files:** 0.00 MB - **Size of the generated dataset:** 219.04 MB - **Total amount of disk used:** 219.04 MB An example of 'train' looks as follows. ``` { "description": "媳妇怀孕五个月了经检查右侧肾积水、过了半月左侧也出现肾积水、她要拿掉孩子、怎么办?", "id": 536714, "label": 8, "question": "孕5个月检查右侧肾积水孩子能要吗?" } ``` ### Data Fields The data fields are the same among all splits. #### age_classification - `question`: a `string` feature. - `description`: a `string` feature. - `label`: a classification label, with possible values including `0-1岁` (0), `1-2岁` (1), `2-3岁` (2). - `id`: a `int32` feature. #### qa - `question`: a `string` feature. - `answer`: a `string` feature. - `id`: a `int32` feature. #### summarization - `description`: a `string` feature. - `question`: a `string` feature. - `id`: a `int32` feature. #### topic_classification - `question`: a `string` feature. - `description`: a `string` feature. - `label`: a classification label, with possible values including `产褥期保健` (0), `儿童过敏` (1), `动作发育` (2), `婴幼保健` (3), `婴幼心理` (4). - `id`: a `int32` feature. ### Data Splits | name |train |validation| test | |--------------------|-----:|---------:|-----:| |age_classification |134852| 19323| 38318| |qa |747888| 106842|213681| |summarization |747888| 106842|213681| |topic_classification|613036| 87519|175363| ## Dataset Creation ### Curation Rationale [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ### Source Data #### Initial Data Collection and Normalization [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) #### Who are the source language producers? [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ### Annotations #### Annotation process [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) #### Who are the annotators? [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ### Personal and Sensitive Information [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ## Considerations for Using the Data ### Social Impact of Dataset [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ### Discussion of Biases [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ### Other Known Limitations [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ## Additional Information ### Dataset Curators [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ### Licensing Information [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) ### Citation Information ``` @inproceedings{xu-etal-2020-matinf, title = "{MATINF}: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization", author = "Xu, Canwen and Pei, Jiaxin and Wu, Hongtao and Liu, Yiyu and Li, Chenliang", booktitle = "Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics", month = jul, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://www.aclweb.org/anthology/2020.acl-main.330", pages = "3586--3596", } ``` ### Contributions Thanks to [@JetRunner](https://github.com/JetRunner) for adding this dataset.

MATINF是第一个联合标注的大规模数据集,适用于分类、问答和摘要生成任务。该数据集包含107万个问答对,带有用户生成的问题描述和人工标注的类别。基于这些丰富的信息,MATINF适用于三大NLP任务:分类、问答和摘要生成。数据集包括四个配置:age_classification(年龄分类)、topic_classification(主题分类)、summarization(摘要生成)和qa(问答)。每个配置都有详细的数据字段和分割信息。

提供机构:
WHUIR
原始信息汇总

数据集卡片 "matinf"

数据集描述

数据集概述

MATINF 是一个联合标注的大型数据集,适用于分类、问答和摘要任务。该数据集包含 107 万个问题-答案对,具有人工标注的类别和用户生成的问题描述。基于这些丰富的信息,MATINF 适用于三大自然语言处理任务,包括分类、问答和摘要。

数据集结构

数据实例

age_classification

  • 大小: 48.39 MB
  • 示例: json { "description": "6个月的时候去儿宝检查,医生说宝宝的分胯动作做的不好,说最好去儿童医院看看,但我家宝宝很好,感觉没有什么不正常啊,请教一下,分胯做的不好,有什么不好吗?", "id": 88016, "label": 0, "question": "医生说宝宝的分胯动作不好" }

qa

  • 大小: 268.69 MB
  • 示例: json { "answer": "我一个同学的孩子就是发现了肾积水,治疗了一段时间,结果还是越来越多,没办法就打掉了。虽然舍不得,但是还是要忍痛割爱,不然以后孩子真的有问题,大人和孩子都受罪。不过,这个最后的决定还要你自己做,毕竟是你的宝宝。", "id": 536714, "question": "孕5个月检查右侧肾积水孩子能要吗?" }

summarization

  • 大小: 258.88 MB
  • 示例: json { "description": "宝宝有中度HIE,但原因未查明,这是他出生后脸上红的几道,嘴唇深红近紫,请问这是像缺氧的表现吗?", "id": 173649, "question": "宝宝脸上红的几道嘴唇深红近紫是像缺氧的表现吗?" }

topic_classification

  • 大小: 219.04 MB
  • 示例: json { "description": "媳妇怀孕五个月了经检查右侧肾积水、过了半月左侧也出现肾积水、她要拿掉孩子、怎么办?", "id": 536714, "label": 8, "question": "孕5个月检查右侧肾积水孩子能要吗?" }

数据字段

age_classification

  • question: 字符串特征。
  • description: 字符串特征。
  • label: 分类标签,可能值包括 0-1岁 (0), 1-2岁 (1), 2-3岁 (2)。
  • id: 整数特征。

qa

  • question: 字符串特征。
  • answer: 字符串特征。
  • id: 整数特征。

summarization

  • description: 字符串特征。
  • question: 字符串特征。
  • id: 整数特征。

topic_classification

  • question: 字符串特征。
  • description: 字符串特征。
  • label: 分类标签,可能值包括 产褥期保健 (0), 儿童过敏 (1), 动作发育 (2), 婴幼保健 (3), 婴幼心理 (4)。
  • id: 整数特征。

数据分割

名称 训练集 验证集 测试集
age_classification 134852 19323 38318
qa 747888 106842 213681
summarization 747888 106842 213681
topic_classification 613036 87519 175363
搜集汇总
数据集介绍
构建方式
MATINF数据集由武汉大学信息检索团队构建,旨在为母婴健康领域的自然语言处理研究提供大规模、多任务标注语料。该数据集从中文互联网母婴社区中采集了超过一百万条用户生成的问答对,涵盖孕期保健、婴幼喂养、儿童疾病等多元主题。每条数据均经过人工标注,形成年龄分类、主题分类、问答匹配与摘要生成四个子任务配置。在年龄分类任务中,问题被划分为0-1岁、1-2岁和2-3岁三个层级;主题分类则细化为18个细粒度类别,如产褥期保健、婴幼心理等。问答与摘要任务共享同一语料来源,分别保留问题-答案对与问题-描述对。数据集按6:2:2比例划分为训练集、验证集和测试集,各子集规模从数万到数十万不等,确保了模型评估的统计可靠性。
使用方法
使用MATINF数据集时,研究者可通过HuggingFace的datasets库便捷加载,支持按任务名称(如'age_classification'、'topic_classification'、'qa'、'summarization')分别调用。在分类任务中,可基于问题与描述字段预测年龄或主题标签;问答任务则利用问题字段检索对应答案,适用于检索式或生成式问答模型的训练;摘要任务以描述为源文本、问题为目标摘要,适合序列到序列模型的微调。数据集已预设训练、验证与测试划分,用户可直接用于模型训练与评估。建议针对不同任务设计专属预处理流水线,例如对分类任务进行标签编码,对摘要任务进行文本截断与填充,以适配各类深度学习框架的输入要求。
背景与挑战
背景概述
MATINF数据集由武汉大学信息检索团队于2020年创建,核心研究人员包括徐灿文、裴嘉欣、吴洪涛、刘一宇和李晨亮,相关成果发表于ACL 2020。该数据集聚焦于母婴健康领域,是首个同时支持分类、问答和摘要三项自然语言处理任务的联合标注大规模数据集。其核心研究问题在于如何利用丰富的用户生成内容,构建一个多任务学习基准,以推动母婴健康信息处理技术的发展。MATINF包含107万对问答数据,覆盖年龄分类、主题分类、问答匹配及摘要生成等任务,为母婴健康领域的智能问答系统、知识图谱构建和对话系统提供了宝贵的数据基础,显著提升了该领域研究的深度与广度。
当前挑战
MATINF所解决的领域问题在于母婴健康信息处理中多任务学习的统一建模挑战,尤其是如何从用户生成的复杂问答中同时提取分类、摘要和问答信息。构建过程中面临的主要挑战包括:1)数据标注的复杂性与一致性,需对107万对问答进行人工标注,涵盖18个主题类别和3个年龄阶段,确保标注质量;2)数据来源的多样性,用户问题涉及大量口语化表达和医学专业术语,增加了自然语言理解的难度;3)多任务联合学习的设计,需平衡分类、问答和摘要任务之间的相关性,避免任务冲突;4)数据隐私保护,母婴健康信息高度敏感,需在公开数据集中去除个人身份信息,确保合规性。
常用场景
经典使用场景
MATINF作为首个同时标注了分类、问答与摘要三种自然语言处理任务的大规模母婴领域数据集,其经典使用场景聚焦于多任务联合建模与基准评测。研究者可利用其丰富的问答对、用户生成的问题描述及人工标注的类别标签,在年龄分类、主题分类、问答生成和文本摘要四个子任务上训练和评估模型。该数据集为母婴健康知识图谱构建、智能问答系统开发以及多任务学习范式探索提供了标准化的数据基础,尤其适合评估模型在低资源场景下的泛化能力和跨任务迁移表现。
解决学术问题
MATINF有效缓解了母婴领域缺乏大规模、高质量、多任务标注语料的学术困境,解决了传统数据集仅支持单一任务且领域覆盖窄的局限。它使得研究者能够在统一的框架下同时探索分类、问答与摘要任务之间的内在关联,推动多任务学习、联合训练和知识共享等前沿方向的发展。该数据集的出现显著提升了母婴健康信息处理的可复现性和可比性,为跨任务迁移学习、弱监督学习以及人机交互中的语义理解研究提供了关键支撑。
实际应用
在实际应用中,MATINF支撑了母婴健康智能助手的核心功能模块,包括基于用户提问的自动化年龄与主题分类、针对育儿困惑的精准问答匹配,以及长文本咨询的智能摘要生成。这些能力可被集成到在线问诊平台、孕期管理APP和育儿社区中,帮助家长快速获取针对性知识,降低信息检索成本。此外,该数据集还可用于构建母婴领域垂直搜索引擎的语义索引,优化医疗资源分配,提升基层健康咨询服务的效率与可及性。
数据集最近研究
最新研究方向
在母婴健康智能问答与文本挖掘的前沿领域,WHUIR/matinf数据集凭借其百万级规模及分类、问答、摘要三大任务的联合标注特性,正推动着多任务学习与迁移学习在垂直医疗场景中的深度应用。当前研究热点聚焦于利用该数据集构建面向孕产期保健与婴幼儿常见病的智能诊断辅助系统,通过融合用户生成的描述性文本与结构化标签,探索细粒度年龄分类、症状主题识别及个性化健康建议生成。结合大语言模型与知识图谱的协同推理,该数据集为低资源场景下的母婴领域自然语言处理提供了高质量基准,其跨任务泛化能力在降低医疗信息不对称、提升基层妇幼保健服务可及性方面展现出显著社会价值。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务