imvladikon/he_cnn_dailymail
收藏资源简介:
--- dataset_info: features: - name: id dtype: string - name: article dtype: string - name: highlights dtype: string - name: article_en dtype: string - name: highlights_en dtype: string splits: - name: train num_bytes: 2945431371 num_examples: 287113 - name: validation num_bytes: 134808274 num_examples: 13368 - name: test num_bytes: 116636491 num_examples: 11490 download_size: 1781960238 dataset_size: 3196876136 task_categories: - summarization language: - he --- # Dataset Card for "he_cnn_dailymail" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息: 特征字段: - 名称:id,数据类型:字符串 - 名称:article,数据类型:字符串 - 名称:highlights(摘要),数据类型:字符串 - 名称:article_en(英文原文文章),数据类型:字符串 - 名称:highlights_en(英文原文摘要),数据类型:字符串 数据集划分: - 划分名称:train(训练集),字节数:2945431371,样本数:287113 - 划分名称:validation(验证集),字节数:134808274,样本数:13368 - 划分名称:test(测试集),字节数:116636491,样本数:11490 下载大小:1781960238 数据集总大小:3196876136 任务类别: - 文本摘要(summarization) 语言: - 希伯来语(he) # 希伯来语版CNN/DailyMail(he_cnn_dailymail)数据集卡片 【需补充更多信息】(https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集概述
数据集名称
- 名称: he_cnn_dailymail
数据集特征
- 特征列表:
- id: 数据类型为字符串
- article: 数据类型为字符串
- highlights: 数据类型为字符串
- article_en: 数据类型为字符串
- highlights_en: 数据类型为字符串
数据集分割
- 训练集:
- 样本数量: 287113
- 数据大小: 2945431371 字节
- 验证集:
- 样本数量: 13368
- 数据大小: 134808274 字节
- 测试集:
- 样本数量: 11490
- 数据大小: 116636491 字节
数据集大小
- 下载大小: 1781960238 字节
- 数据集总大小: 3196876136 字节
任务类别
- 任务: 摘要生成
语言
- 语言: 希伯来语



