aixk/vlite3.1-coder-dataset
收藏官方服务:
资源简介:
该数据集是一个用于自然语言处理(NLP)任务的数据集,包含训练数据,特征包括input_ids(表示输入文本的token ID序列,int32列表)、attention_mask(表示注意力掩码,int8列表)和labels(表示标签序列,int64列表)。数据集共有645,522个训练示例,文件大小约为4.3GB,下载大小约为5.3GB。
This dataset is designed for natural language processing (NLP) tasks, containing training data with features including input_ids (a list of int32 representing token IDs for input text), attention_mask (a list of int8 representing attention masks), and labels (a list of int64 representing label sequences). The dataset consists of 645,522 training examples, with a file size of approximately 4.3GB and a download size of approximately 5.3GB.
提供机构:
aixk


