遇见数据集

BlackwoodAI/pretrain-shards

收藏
Hugging Face2026-05-12 更新2026-05-31 收录
官方服务:

资源简介:

该数据集名为Pretrain shards,是一个用于预训练的预分词分片集合。它包含三个子目录:code(代码)、text(文本)和math(数学),每个子目录下包含多个分片文件,格式为.bin(存储uint32小端序的token ID)和.json(包含每个分片的文档长度和统计信息)。数据集基于CC BY-NC 4.0许可证,仅限非商业用途,主要用于内部研究。用户可以通过HuggingFace Hub下载特定分片,例如代码分片。

The dataset named Pretrain shards is a collection of pre-tokenized shards for pretraining. It includes three subdirectories: code, text, and math, each containing multiple shard files in .bin format (storing uint32 little-endian token IDs) and .json format (containing per-shard document lengths and statistics). Licensed under CC BY-NC 4.0 for non-commercial use only, it is intended for internal research purposes. Users can download specific shards, such as code shards, via the HuggingFace Hub.

提供机构:
BlackwoodAI
二维码
社区交流群
二维码
科研交流群
商业服务