遇见数据集

kothasuhas/dclm_tokenized_n307888_ctx4096_pg_sp4096_eos

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个用于自然语言处理任务的大规模文本数据集,包含307,888个训练示例,每个示例由长度为4096的input_ids序列组成,序列类型为uint16,表示文本的token化表示。数据集总大小约为2.52 GB,下载大小约为1.88 GB,适用于训练语言模型或文本生成任务。

This dataset is a large-scale text dataset for natural language processing tasks, containing 307,888 training examples. Each example consists of an input_ids sequence with a length of 4096, of type uint16, representing tokenized text. The total dataset size is approximately 2.52 GB, with a download size of about 1.88 GB, suitable for training language models or text generation tasks.

提供机构:
kothasuhas
二维码
社区交流群
二维码
科研交流群
商业服务