Estwld/atomic2020-instruct-drop_duplicates
收藏资源简介:
--- dataset_info: features: - name: knowledge_type dtype: string - name: task_type dtype: string - name: input dtype: string - name: output dtype: string splits: - name: train num_bytes: 755049360 num_examples: 2016508 - name: validation num_bytes: 70762021 num_examples: 189228 - name: test num_bytes: 108704624 num_examples: 287472 download_size: 79239881 dataset_size: 934516005 --- # Dataset Card for "atomic2020-instruct-drop_duplicates" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息: 特征字段: - 字段名:knowledge_type(知识类型),数据类型:字符串 - 字段名:task_type(任务类型),数据类型:字符串 - 字段名:input(输入),数据类型:字符串 - 字段名:output(输出),数据类型:字符串 数据拆分: - 拆分名称:train(训练集),字节数:755049360,样本数:2016508 - 拆分名称:validation(验证集),字节数:70762021,样本数:189228 - 拆分名称:test(测试集),字节数:108704624,样本数:287472 下载大小:79239881,数据集总大小:934516005 --- # 「atomic2020-instruct-drop_duplicates」数据集卡片 [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集概述
特征信息
- knowledge_type: 数据类型为字符串。
- task_type: 数据类型为字符串。
- input: 数据类型为字符串。
- output: 数据类型为字符串。
数据分割
- train:
- 字节数: 755049360
- 样本数: 2016508
- validation:
- 字节数: 70762021
- 样本数: 189228
- test:
- 字节数: 108704624
- 样本数: 287472
数据集大小
- 下载大小: 79239881 字节
- 数据集大小: 934516005 字节




