遇见数据集

mrm8488/large_spanish_corpus_ds_tokenized_and_gropuped

收藏
Hugging Face2023-02-13 更新2024-03-04 收录
官方服务:

资源简介:

--- dataset_info: features: - name: input_ids sequence: int32 splits: - name: train num_bytes: 16824296700 num_examples: 4103487 - name: test num_bytes: 885489300 num_examples: 215973 download_size: 8311975924 dataset_size: 17709786000 --- # Dataset Card for "large_spanish_corpus_ds_tokenized_and_gropuped" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)

数据集信息: 特征字段: - 字段名:输入词元ID(input_ids),数据格式为int32序列 数据集拆分: - 拆分名称:训练集(train),字节数:16824296700,样本数量:4103487 - 拆分名称:测试集(test),字节数:885489300,样本数量:215973 下载大小:8311975924 数据集总大小:17709786000 --- # 「large_spanish_corpus_ds_tokenized_and_gropuped」数据集卡片 [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)

提供机构:
mrm8488
原始信息汇总

数据集概述

数据集名称

large_spanish_corpus_ds_tokenized_and_gropuped

数据集特征

  • 名称: input_ids
  • 序列类型: int32

数据集分割

  • 训练集
    • 样本数量: 4103487
    • 数据大小: 16824296700 字节
  • 测试集
    • 样本数量: 215973
    • 数据大小: 885489300 字节

数据集大小

  • 下载大小: 8311975924 字节
  • 总数据大小: 17709786000 字节
二维码
社区交流群
二维码
科研交流群
商业服务