TR-HASH-Pretraining-125B-Agentic-32K
收藏资源简介:
TR-HASH Pretraining 125B Agentic 32K 是一个私有、来源策划的预训练数据集,专为 TR-HASH Agentic 32K 模型系列设计。数据集包含总计 125B 个打包 token 曝光,其中 75B 来自 Foundation 桶(涵盖英语和法语知识、教育网页文本、数学以及合成教科书),50B 来自 Agentic 桶(涵盖教育代码、程序、调试、验证、规划、数学推理以及受限的工具使用轨迹)。数据集使用经过验证的 32,000-ID 版本的 AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic,与旧版 tokenizer 不兼容。该数据集采用高吞吐量直接构建方式,从策划的上游子集读取文档并直接 token 化,但未进行每文档质量过滤、代理信号过滤、基准去污染或全局精确文档去重。训练课程分为两个阶段:首先 foundation-first(60B foundation + 15B agentic),然后 agentic-intensification(15B foundation + 35B agentic)。数据集保持私有,直到源许可证、分片哈希、token 预算和此物化披露经过审计。所有源仓库、配置、修订、token 预算和许可证审计说明均固定在发布的 _metadata/config.json 中。
TR-HASH Pretraining 125B Agentic 32K is a private, source-curated pretraining dataset designed for the TR-HASH Agentic 32K model series. The dataset contains a total of 125B packed token exposures, with 75B from the Foundation bucket (covering English and French knowledge, educational web text, mathematics, and synthetic textbooks) and 50B from the Agentic bucket (covering educational code, programs, debugging, verification, planning, mathematical reasoning, and constrained tool use trajectories). The dataset uses a verified 32,000-ID version of AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic, which is incompatible with the older tokenizer. It is constructed through a high-throughput direct approach, reading documents from curated upstream subsets and directly tokenizing them, without per-document quality filtering, agent signal filtering, benchmark decontamination, or global exact document deduplication. The training curriculum consists of two stages: first foundation-first (60B foundation + 15B agentic), then agentic-intensification (15B foundation + 35B agentic). The dataset remains private until the source licenses, shard hashes, token budgets, and this materialization disclosure are audited. All source repositories, configurations, revisions, token budgets, and license audit notes are fixed in the published _metadata/config.json.
TR-HASH Pretraining 125B Agentic 32K 数据集详情
概述
这是一个私有、来源精选的预训练数据集,专为 TR-HASH Agentic 32K 模型系列设计。数据集包含 1250亿个打包令牌暴露,其中基础内容 750亿,智能体/程序性内容 500亿。
主要特征
- 语言:英语、法语
- 任务类型:文本生成
- 许可证:其他(专有许可)
- 分词器:使用不可变的、经过验证的 32,000 词汇量修订版
AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic,与旧版 TR-HASH 32K 分词器不兼容
数据集构成
| 数据桶 | 令牌数 | 用途 |
|---|---|---|
| 基础 | 750亿 | 英语和法语知识、教育网页文本、数学和合成教科书 |
| 智能体 | 500亿 | 教育代码、程序、调试、验证、规划、数学推理和限额工具使用轨迹 |
所有源仓库、配置、修订版本、令牌预算和许可证审计说明均固定在发布的 _metadata/config.json 文件中。
构建方式说明
- 采用来源精选直接构建的高吞吐量方式,从精选的上游子集中读取文本并直接进行分词
- 明确不声称包含:逐文档质量过滤、智能体信号过滤、基准去污或全局精确文档去重
- 每个10亿令牌的
uint16分片在上传后,会校验远程大小和 SHA-256,提交到重启状态,然后才在本地删除 - 数据集在源许可证、分片哈希、令牌预算和物化披露完成审计前保持私有状态
课程安排
运行时计划将每个打包行消费一次:
- 基础优先阶段:600亿基础 + 150亿智能体
- 智能体强化阶段:150亿基础 + 350亿智能体
分片边界与阶段划分对齐。unique_tokens 和 trained_tokens 描述打包令牌位置;由于直接构建中禁用了全局去重,源文档可能重复出现。





