Arithmetic-1.5M
收藏资源简介:
ArithMark 2.0预训练数据集是一个专门为参数规模小于1000万的语言模型设计的预训练数据集。其核心目标是提升这些小型模型在ArithMark-2.0算术基准测试上的性能,使其表现显著超越随机猜测水平。数据集包含两个版本的文件:`arithmark_raw.parquet`文件保留了数据的原始顺序,而`arithmark_curriculum.parquet`文件则按照课程学习的顺序进行了编排,推荐使用此版本进行模型训练。该数据集于2026年6月5日生成。
The ArithMark 2.0 Pre-training Dataset is a specialized dataset tailored for language models with fewer than 10 million parameters. Its core goal is to improve the performance of these small-scale language models on the ArithMark-2.0 arithmetic benchmark, ensuring their results significantly outperform random guessing levels. The dataset includes two file versions: the `arithmark_raw.parquet` file retains the original order of the data, while the `arithmark_curriculum.parquet` file is arranged in a curriculum learning sequence, and this version is recommended for model training. This dataset was generated on June 5, 2026.
数据集概述:ArithMark 2.0 预训练数据集
- 语言:英语(en)
- 标签:算术(arithmetic)、预训练(pretraining)、低于1000万参数(sub-10M)、ArithMark
- 用途:专为针对ArithMark-2.0基准测试、期望达到高于随机水平性能的低于1000万参数大语言模型生成的预训练数据集。
- 文件内容:
arithmark_raw.parquet:原始顺序数据。arithmark_curriculum.parquet:课程学习排序数据(建议用于训练)。
- 生成时间:2026-06-05 15:57 UTC





