cmu-lti/osim-mid-training
收藏资源简介:
该数据集是ODYSSIM行为基础模型实验中使用的中期训练语料库。数据以按数据集/来源分组的parquet分片形式存储,包含训练分片和对应来源的保留测试文件。该版本将之前暂存的数据集Xuhui/sft_processed_large_split_v3转移到CMU-LTI组织用于论文发布。统计摘要显示:2140万训练行,约106亿估计词元,166个存储库文件总计约21GB。语料库聚合了多个上游数据集,具有不同的许可证和使用条款,建议在重新分发或商业使用前查阅ODYSSIM论文中的来源级文档和引用。
This repository contains the midtraining corpus used for ODYSSIM behavioral foundation model experiments. The corpus is stored as parquet shards grouped by dataset/source, with train shards and held-out test files for the corresponding sources. The release mirrors the previously staged dataset `Xuhui/sft_processed_large_split_v3` into the CMU-LTI organization for the paper release. Summary statistics from the audit pass: 21.4M train rows, approximately 10.6B estimated tokens, and 166 repository files totaling about 21GB on Hugging Face. The corpus aggregates many upstream datasets with different licenses and use terms. Please consult the source-level documentation and citations in the ODYSSIM paper before redistribution or commercial use.




