Beetle-Data/en-for-is-100M-pretok
收藏官方服务:
资源简介:
这是一个名为Beetle-Data/en-for-is-100M-pretok的预分词数据集,包含100M token规模的英语文本,用于NLP任务。数据被预处理为513个token的打包序列块,确保没有跨文档混合,以提高训练效率。数据集以分片形式存储和上传,使用Parquet文件格式,并包含一个标记文件来指示上传完成状态。
This is a pretokenized dataset named Beetle-Data/en-for-is-100M-pretok, containing 100M tokens of English text for NLP tasks. The data is preprocessed into packed sequences of 513-token chunks, with no cross-document bleeding to enhance training efficiency. The dataset is stored and uploaded in sharded Parquet files, with a marker file indicating completion of upload.
提供机构:
Beetle-Data


