Platinum Benchmarks
收藏资源简介:
Platinum Benchmarks是一组经过精心策划的测试集,目的是最小化标签错误和歧义,以便能够评估语言模型在任务上是否能够达到100%的准确度。这些测试集覆盖了数学、逻辑、表格理解、阅读理解、常识推理和视觉理解等多个能力类别,包含的问题从简单的单一操作到高年级的数学问题不等。作者通过对现有十五个流行测试集的修订,移除或纠正了错误和歧义,从而构建了这些Platinum Benchmarks。这些测试集可用于评估前沿语言模型在不同难度级别任务上的可靠性边界。
Platinum Benchmarks is a carefully curated collection of test sets designed to minimize labeling errors and ambiguity, enabling the evaluation of whether language models can achieve 100% accuracy on target tasks. These test sets cover multiple capability categories including mathematics, logic, table understanding, reading comprehension, commonsense reasoning and visual understanding, with problems ranging from simple single-step operations to advanced upper-level mathematical problems. The authors constructed these Platinum Benchmarks by revising fifteen existing popular test sets, removing or correcting labeling errors and ambiguities. These test sets can be used to evaluate the reliability boundaries of state-of-the-art language models on tasks across different difficulty levels.




