官方服务:
资源简介:
WAIS-III tasks.
应用场景:
创建时间:
2020-08-21
相关数据集
Platinum Benchmarks
Platinum Benchmarks是一组经过精心策划的测试集,目的是最小化标签错误和歧义,以便能够评估语言模型在任务上是否能够达到100%的准确度。这些测试集覆盖了数学、逻辑、表格理解、阅读理解、常识推理和视觉理解等多个能力类别,包含的问题从简单的单一操作到高年级的数学问题不等。作者通过对现有十五个流行测试集的修订,移除或纠正了错误和歧义,从而构建了这些Platinum Benchmarks。
arXiv2025-02-06 更新190
Simple Multimodal Algorithmic Reasoning Task Dataset (SMART-101)
Introduction Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating ar
NIAID Data Ecosystem110
Performance across all conditions for the crows across both apparatuses.
Results reflect results of Wilcoxon 1-sample signed ranks tests–chance value = 50%. Significant p-values (< .05) highlighted in bold. S1-8 stands for the number of sessions in this condition.
NIAID Data Ecosystem60
maywell/LogicKor
LogicKor是一个用于评估韩国语语言模型在多个领域思考能力的多轮基准测试数据集。该数据集包含6个类别(推理、数学、写作、编码、理解、语法)的42个多轮提示,旨在通过LLM-as-a-judge方式测量模型在不同领域的表现。每个类别都有具体的描述,例如推理涉及逻辑思维和问题解决,数学涉及数学概念和计算等。
Hugging Face2024-06-09 更新130
Repeated Testing Study of Cognitive Ability Tests and Working Memory Capacity
This dataset contains data of N = 221 participants, that took part in a repeated testing study of cognitive ability tests. Participants were aged 18-77 of various backgrounnds, although the majority w
CESSDA2023-11-16 更新100



