遇见数据集

MLP-SEMO/CDT_qa_datasets_

收藏
Hugging Face2024-07-14 更新2024-07-22 收录
官方服务:

资源简介:

该数据集包含两个配置版本:1.9M和7.6M。每个配置都包含上下文和响应两个特征,数据类型均为字符串。1.9M配置的训练集包含1,943,744个样本,占用5,991,496,719字节,下载大小为3,678,970,771字节。7.6M配置的训练集包含7,637,632个样本,占用23,544,145,772字节,下载大小为14,459,063,097字节。数据文件路径分别为1.9M/train-*和7.6M/train-*。

The dataset includes two configurations: 1.9M and 7.6M. Each configuration contains two features: context and response, both of which are of string type. The 1.9M configurations training set consists of 1,943,744 samples, occupying 5,991,496,719 bytes, with a download size of 3,678,970,771 bytes. The 7.6M configurations training set consists of 7,637,632 samples, occupying 23,544,145,772 bytes, with a download size of 14,459,063,097 bytes. The data file paths are 1.9M/train-* and 7.6M/train-* respectively.

提供机构:
MLP-SEMO
二维码
社区交流群
二维码
科研交流群
商业服务