mrm8488/large_spanish_corpus_ds_tokenized_and_gropuped
收藏资源简介:
--- dataset_info: features: - name: input_ids sequence: int32 splits: - name: train num_bytes: 16824296700 num_examples: 4103487 - name: test num_bytes: 885489300 num_examples: 215973 download_size: 8311975924 dataset_size: 17709786000 --- # Dataset Card for "large_spanish_corpus_ds_tokenized_and_gropuped" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息: 特征字段: - 字段名:输入词元ID(input_ids),数据格式为int32序列 数据集拆分: - 拆分名称:训练集(train),字节数:16824296700,样本数量:4103487 - 拆分名称:测试集(test),字节数:885489300,样本数量:215973 下载大小:8311975924 数据集总大小:17709786000 --- # 「large_spanish_corpus_ds_tokenized_and_gropuped」数据集卡片 [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集概述
数据集名称
large_spanish_corpus_ds_tokenized_and_gropuped
数据集特征
- 名称: input_ids
- 序列类型: int32
数据集分割
- 训练集
- 样本数量: 4103487
- 数据大小: 16824296700 字节
- 测试集
- 样本数量: 215973
- 数据大小: 885489300 字节
数据集大小
- 下载大小: 8311975924 字节
- 总数据大小: 17709786000 字节



