遇见数据集

Beetle-Data/en-for-he-100M-pretok

收藏
Hugging Face2026-05-16 更新2026-05-31 收录
官方服务:

资源简介:

Beetle-Data/en-for-he-100M-pretok数据集是一个预分块的文本数据集,包含约1亿个标记的英语到希伯来语(或相关语言)的预训练数据。数据以513个标记的打包序列形式组织,确保没有跨文档混合,并采用分片增量上传方式,上传完成后通过data/_finalized.json文件标记为最终版本。

The Beetle-Data/en-for-he-100M-pretok dataset is a pretokenized text dataset containing approximately 100 million tokens for English-to-Hebrew (or related language) pretraining. The data is organized into packed sequences of 513 tokens each, with no cross-document bleeding, and is sharded incrementally, with a marker file data/_finalized.json committed once all parts are uploaded.

提供机构:
Beetle-Data
二维码
社区交流群
二维码
科研交流群
商业服务