yardeny/processed_t5_small_context_len_512
收藏官方服务:
资源简介:
--- dataset_info: features: - name: input_ids sequence: int32 - name: attention_mask sequence: int8 splits: - name: train num_bytes: 17763456912.0 num_examples: 6917234 download_size: 6975491955 dataset_size: 17763456912.0 --- # Dataset Card for "processed_t5_small_context_len_512" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息: 特征: - 输入ID(input_ids):类型为int32的序列 - 注意力掩码(attention_mask):类型为int8的序列 数据集划分: - 训练集(train):字节大小为17763456912.0,样本量为6917234 下载大小:6975491955 数据集总大小:17763456912.0 # 「processed_t5_small_context_len_512」数据集卡片 [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
提供机构:
yardeny原始信息汇总
数据集概述
数据特征
- input_ids: 序列类型为 int32
- attention_mask: 序列类型为 int8
数据分割
- train: 包含 6917234 个样本,总大小为 17763456912.0 字节
数据大小
- 下载大小: 6975491955 字节
- 数据集大小: 17763456912.0 字节



