遇见数据集

CLIWorks/Spider-FLEXITOKENS-FP8

收藏
Hugging Face2026-05-10 更新2026-05-31 收录
官方服务:

资源简介:

FineWeb-Edu字节级数据集是一个预分词的文本数据集,包含约170亿个令牌,以字节级别编码。数据集包含14个分片,每个分片约2.5GB,包含约12亿个令牌。数据格式为无标头的uint16 .bin文件,使用UTF-8字节编码(0-255)和特殊令牌:BOS(开始符)=257,EOS(结束符)=258,PAD(填充符)=256。数据集用于训练Spider-FLEXITOKENS模型,支持FP8训练流程。

The FineWeb-Edu Byte-Level Dataset is a pre-tokenized text dataset containing approximately 17 billion tokens at the byte level. It consists of 14 shards, each approximately 2.5GB in size and containing about 1.2 billion tokens. The data is stored in headerless uint16 .bin files, using UTF-8 byte encoding (0-255) with special tokens: BOS=257, EOS=258, PAD=256. This dataset is used for training the Spider-FLEXITOKENS model and supports FP8 training pipelines.

提供机构:
CLIWorks
二维码
社区交流群
二维码
科研交流群
商业服务