相关数据集
Labagaite/fr-summarizer-dataset
--- dataset_info: features: - name: fr-summarizer-dataset dtype: string - name: content dtype: string splits: - name: train num_bytes: 13739369 num_examples: 1968 - name: v
Hugging Face2024-04-13 更新190
Multi-XScience
该数据集名为Multi-XScience,是一个大规模的多文档摘要数据集,它由科学论文构建而成,专注于根据论文摘要及其引用的文章编写论文的相关工作部分。该数据集旨在最大化对抽象模型的实用性,并包含高水平的抽象性,这使得它对于摘要模型来说颇具挑战性。它是一个大规模的数据集,所涉及的任务是多文档摘要。
arXiv130
arthurmluz/GPTextSum_data-temario_results
--- dataset_info: features: - name: id dtype: int64 - name: text dtype: string - name: summary dtype: string - name: gen_summary dtype: string - name: rouge struct:
Hugging Face2023-11-15 更新120
The Pile Dataset
The Pile is a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets combined together. Datasheet: Datasheet for the Pile
paperswithcode.com190
autoevaluate/autoeval-staging-eval-project-84760c85-7314786
该数据集包含由AutoTrain生成的模型预测结果,用于文本摘要任务。使用的模型是philschmid/distilbart-cnn-12-6-samsum,数据集为samsum。
Hugging Face2022-06-24 更新120



