imminsik/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_gpt2_1757511792
收藏数据链接:
官方服务:
资源简介:
这是一个用于OpenAI的Summarize from Feedback任务的Reddit TL;DR数据集。数据集包含了帖子的唯一标识符、所属subreddit、标题、正文、摘要以及参考回复等字段。此外,还包括了一些预处理后添加的字段,如长度受限的查询、查询的标记化版本、参考回复的标记化版本等。数据集分为训练集、验证集和测试集,适用于文本摘要任务。
This is the Reddit TL;DR dataset for OpenAIs Summarize from Feedback task. The dataset includes fields such as unique identifier for the post, subreddit, title, body of the post, summary, and reference response. Additionally, it contains fields added by the preprocessing script, such as length-limited query, tokenized version of the query, tokenized version of the reference response, etc. The dataset is split into training, validation, and test sets, suitable for text summarization tasks.
提供机构:
imminsik


