遇见数据集

timaeus/prefix-ab

收藏
Hugging Face2024-06-22 更新2024-06-29 收录
官方服务:

资源简介:

该数据集包含文本和掩码序列两个主要特征,主要用于训练模型。数据集包含5000个训练样本,总大小为386033字节,下载大小为145563字节。数据文件路径在默认配置下指定为train-*。

This dataset includes two main features: text and mask sequences, primarily used for training models. The dataset contains 5000 training samples with a total size of 386033 bytes and a download size of 145563 bytes. The data file path is specified as train-* under the default configuration.

提供机构:
timaeus
原始信息汇总

数据集概述

数据集信息

  • 特征

    • text:数据类型为字符串。
    • mask:数据类型为整数序列。
  • 分割

    • train:包含5000个样本,占用386033.0字节。
  • 下载大小:145563字节。

  • 数据集大小:386033.0字节。

配置

  • 配置名称:default
    • 数据文件
      • train:路径为data/train-*
二维码
社区交流群
二维码
科研交流群
商业服务