antolin/tlc_interduplication
收藏官方服务:
资源简介:
--- dataset_info: features: - name: id_within_dataset dtype: int64 - name: snippet dtype: string - name: tokens sequence: string - name: nl dtype: string - name: split_within_dataset dtype: string - name: is_duplicated dtype: bool splits: - name: train num_bytes: 70652063.18677872 num_examples: 53327 - name: test num_bytes: 8799876.304434607 num_examples: 6642 - name: valid num_bytes: 8831673.508786675 num_examples: 6666 download_size: 33772946 dataset_size: 88283613.00000001 --- # Dataset Card for "tlc_interduplication" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
提供机构:
antolin原始信息汇总
数据集概述
数据集特征
- id_within_dataset: 数据类型为
int64 - snippet: 数据类型为
string - tokens: 序列类型为
string - nl: 数据类型为
string - split_within_dataset: 数据类型为
string - is_duplicated: 数据类型为
bool
数据集分割
- train: 字节数为
70652063.18677872,样本数为53327 - test: 字节数为
8799876.304434607,样本数为6642 - valid: 字节数为
8831673.508786675,样本数为6666
数据集大小
- 下载大小:
33772946字节 - 数据集大小:
88283613.00000001字节



