遇见数据集

Dataset For Generating Tl;Dr

收藏
Zenodo2020-09-19 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This is the dataset for the TL;DR challenge containing posts from the Reddit corpus, suitable for abstractive summarization using deep learning. The format is a json file where each line is a JSON object representing a post. The schema of each post is shown below: author: string (nullable = true) body: string (nullable = true) normalizedBody: string (nullable = true) content: string (nullable = true) content_len: long (nullable = true) summary: string (nullable = true) summary_len: long (nullable = true) id: string (nullable = true) subreddit: string (nullable = true) subreddit_id: string (nullable = true) title: string (nullable = true) Specifically, the <strong>content</strong> and <strong>summary</strong> fields can be directly used as inputs to a deep learning model (e.g. Sequence to Sequence model ). The dataset consists of 3,084,410 posts with an average length of 211 words for content, and 25 words for the summary. <strong>Note : </strong>As this is the complete dataset for the challenge, it is up to the participants to split it into training and validation sets accordingly.

提供机构:
Zenodo
创建时间:
2018-02-08
二维码
社区交流群
二维码
科研交流群
商业服务