Aunsiels/InfantBooks
收藏资源简介:
--- annotations_creators: - no-annotation language: - en language_creators: - crowdsourced license: - gpl multilinguality: - monolingual pretty_name: InfantBooks size_categories: - 1M<n<10M source_datasets: - original tags: - research paper - kids - children - books task_categories: - text-generation task_ids: - language-modeling --- # Dataset Card for InfantBooks ## Table of Contents - [Table of Contents](#table-of-contents) - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Languages](#languages) - [Dataset Structure](#dataset-structure) - [Data Instances](#data-instances) - [Data Fields](#data-fields) - [Dataset Creation](#dataset-creation) - [Curation Rationale](#curation-rationale) - [Source Data](#source-data) - [Additional Information](#additional-information) - [Citation Information](#citation-information) ## Dataset Description - **Homepage:** [https://www.mpi-inf.mpg.de/children-texts-for-commonsense](https://www.mpi-inf.mpg.de/children-texts-for-commonsense) - **Paper:** Do Children Texts Hold The Key To Commonsense Knowledge? ### Dataset Summary A dataset of infants/children's books. ### Languages All the books are in English; ## Dataset Structure ### Data Instances malis-friend_BookDash-FKB.txt,"Then a taxi driver, hooting around the yard with his wire car. Mali enjoys playing by himself..." ### Data Fields - title: The title of the book - content: The content of the book ## Dataset Creation ### Curation Rationale The goal of the dataset is to study infant books, which are supposed to be easier to understand than normal texts. In particular, the original goal was to study if these texts contain more commonsense knowledge. ### Source Data #### Initial Data Collection and Normalization We automatically collected kids' books on the web. #### Who are the source language producers? Native speakers. ### Citation Information ``` Romero, J., & Razniewski, S. (2022). Do Children Texts Hold The Key To Commonsense Knowledge? In Proceedings of the 2022 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning. ```
annotations_creators: - 无注释 language: - 英语 language_creators: - 众包 license: - GNU通用公共许可证(GPL) multilinguality: - 单语言 pretty_name: InfantBooks size_categories: - 100万 < 样本数 < 1000万 source_datasets: - 原创数据集 tags: - 研究论文 - 儿童读物 - 儿童 - 书籍 task_categories: - 文本生成 task_ids: - 语言建模 # InfantBooks 数据集卡片 ## 目录 - [目录](#目录) - [数据集描述](#数据集描述) - [数据集摘要](#数据集摘要) - [语言](#语言) - [数据集结构](#数据集结构) - [数据实例](#数据实例) - [数据字段](#数据字段) - [数据集构建](#数据集构建) - [构建依据](#构建依据) - [源数据](#源数据) - [附加信息](#附加信息) - [引用信息](#引用信息) ## 数据集描述 - **主页:** [https://www.mpi-inf.mpg.de/children-texts-for-commonsense] - **论文:** 《儿童文本是否蕴含常识知识的关键?》(Do Children Texts Hold The Key To Commonsense Knowledge?) ### 数据集摘要 本数据集收录婴幼儿及儿童读物。 ### 语言 所有书籍均采用英语撰写。 ## 数据集结构 ### 数据实例 `malis-friend_BookDash-FKB.txt`,"彼时一名出租车司机正驾驶着他的铁丝小车在院子里鸣笛。马里享受独自玩耍的时光……" ### 数据字段 - **标题:** 书籍的标题 - **内容:** 书籍的正文内容 ## 数据集构建 ### 构建依据 本数据集的构建目标为研究婴幼儿读物——相较于普通文本,这类读物理应更易于理解。本项目的初始研究目标尤为聚焦于探究这类文本是否蕴含更为丰富的常识知识。 ### 源数据 #### 初始数据收集与归一化 我们通过自动化手段从网络上收集儿童读物。 #### 源语言创作者是谁? 以英语为母语的使用者。 ### 引用信息 Romero, J. 与 Razniewski, S. (2022). 《儿童文本是否蕴含常识知识的关键?》(Do Children Texts Hold The Key To Commonsense Knowledge?) 收录于2022年自然语言处理经验方法会议与计算自然语言学习联合会议论文集。
数据集概述
数据集描述
数据集总结
- 名称: InfantBooks
- 内容: 包含婴幼儿/儿童书籍的数据集。
语言
- 语言: 英语
- 语言创建者: 众包
数据集结构
数据实例
- 示例:
malis-friend_BookDash-FKB.txt包含书籍内容。
数据字段
- 标题: 书籍的标题
- 内容: 书籍的内容
数据集创建
采集理由
- 目的: 研究婴幼儿书籍,这些书籍被认为比普通文本更易于理解,特别是研究这些文本是否包含更多常识知识。
源数据
- 初始数据收集: 自动从网络上收集儿童书籍
- 源语言生产者: 母语者
附加信息
引用信息
- 作者: Romero, J., & Razniewski, S.
- 出版物: 在2022年联合会议的实证方法在自然语言处理和计算自然语言学习会议论文集
- 论文标题: Do Children Texts Hold The Key To Commonsense Knowledge?




