kothasuhas/dclm_tokenized_n307888_ctx4096_pg_sp4096_eos
收藏数据链接:
官方服务:
资源简介:
该数据集是一个用于自然语言处理任务的大规模文本数据集,包含307,888个训练示例,每个示例由长度为4096的input_ids序列组成,序列类型为uint16,表示文本的token化表示。数据集总大小约为2.52 GB,下载大小约为1.88 GB,适用于训练语言模型或文本生成任务。
This dataset is a large-scale text dataset for natural language processing tasks, containing 307,888 training examples. Each example consists of an input_ids sequence with a length of 4096, of type uint16, representing tokenized text. The total dataset size is approximately 2.52 GB, with a download size of about 1.88 GB, suitable for training language models or text generation tasks.
提供机构:
kothasuhas


