AxiomicLabs/ArithMark-2.0
收藏官方服务:
资源简介:
ArithMark 2.0 是一个程序生成的基准测试,用于评估语言模型的整数算术能力。每个测试项目被格式化为延续式多项选择题:模型会看到一个以等号(=)结尾的算术表达式,必须为正确的数字延续分配最高的可能性分数。该基准设计用于基础模型的对数似然评分,不需要指令遵循、思维链或生成解释,随机猜测的正确率为25%。数据集包含2,500个示例,答案标签完全平衡,涵盖不同难度级别和运算符数量,主题包括加法、减法、混合运算、括号运算等。
ArithMark 2.0 is a procedurally generated benchmark for evaluating integer arithmetic ability in language models. Each item is formatted as a continuation-style multiple-choice problem: the model sees an arithmetic expression ending in = and must assign the highest likelihood to the correct numeric continuation. The benchmark is designed for base-model log-likelihood scoring. It does not require instruction following, chain-of-thought, or generated explanations. Random chance is 25%.
提供机构:
AxiomicLabs


