tyzhu/lmind_nq_train6000_eval6489_v1_doc_qa_first_permute
收藏资源简介:
--- dataset_info: features: - name: answers struct: - name: answer_start sequence: 'null' - name: text sequence: string - name: inputs dtype: string - name: targets dtype: string splits: - name: all_docs_eval num_bytes: 7125701 num_examples: 10925 - name: validation num_bytes: 752802 num_examples: 6489 - name: train_qa num_bytes: 697367 num_examples: 6000 - name: train_ic_qa num_bytes: 4540536 num_examples: 6000 - name: train num_bytes: 38313328 num_examples: 63692 - name: eval_ic_qa num_bytes: 4906186 num_examples: 6489 - name: eval_recite_qa num_bytes: 4912675 num_examples: 6489 - name: all_docs num_bytes: 7126313 num_examples: 10925 - name: eval_qa num_bytes: 752802 num_examples: 6489 - name: train_recite_qa num_bytes: 4546536 num_examples: 6000 - name: first_permute_docs num_bytes: 37615961 num_examples: 57692 download_size: 37058553 dataset_size: 111290207 --- # Dataset Card for "lmind_nq_train6000_eval6489_v1_doc_qa_first_permute" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息: 特征字段: - 名称:answers 结构体: - 名称:answer_start 序列:空值(null) - 名称:text 序列:字符串(string) - 名称:inputs 数据类型(dtype):字符串 - 名称:targets 数据类型(dtype):字符串 数据集划分: - 划分名称:all_docs_eval(全文档评估集) 字节数:7125701 样本数:10925 - 划分名称:validation(验证集) 字节数:752802 样本数:6489 - 划分名称:train_qa(问答训练子集) 字节数:697367 样本数:6000 - 划分名称:train_ic_qa(上下文问答训练子集) 字节数:4540536 样本数:6000 - 划分名称:train(总训练集) 字节数:38313328 样本数:63692 - 划分名称:eval_ic_qa(上下文问答评估子集) 字节数:4906186 样本数:6489 - 划分名称:eval_recite_qa(记忆问答评估子集) 字节数:4912675 样本数:6489 - 划分名称:all_docs(全文档集) 字节数:7126313 样本数:10925 - 划分名称:eval_qa(问答评估子集) 字节数:752802 样本数:6489 - 划分名称:train_recite_qa(记忆问答训练子集) 字节数:4546536 样本数:6000 - 划分名称:first_permute_docs(首次重排文档集) 字节数:37615961 样本数:57692 下载大小:37058553 总数据集大小:111290207 --- # 数据集卡片:"lmind_nq_train6000_eval6489_v1_doc_qa_first_permute" [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集概述
特征信息
- answers: 包含两个子结构
- answer_start: 序列类型为 null
- text: 序列类型为 string
- inputs: 数据类型为 string
- targets: 数据类型为 string
数据分割
- all_docs_eval: 字节数为 7125701,示例数为 10925
- validation: 字节数为 752802,示例数为 6489
- train_qa: 字节数为 697367,示例数为 6000
- train_ic_qa: 字节数为 4540536,示例数为 6000
- train: 字节数为 38313328,示例数为 63692
- eval_ic_qa: 字节数为 4906186,示例数为 6489
- eval_recite_qa: 字节数为 4912675,示例数为 6489
- all_docs: 字节数为 7126313,示例数为 10925
- eval_qa: 字节数为 752802,示例数为 6489
- train_recite_qa: 字节数为 4546536,示例数为 6000
- first_permute_docs: 字节数为 37615961,示例数为 57692
数据集大小
- 下载大小: 37058553 字节
- 数据集大小: 111290207 字节



