NTSEBENCH
收藏资源简介:
NTSEBENCH是由印度理工学院古瓦哈提分校、犹他大学和宾夕法尼亚大学联合创建的一个用于评估大型深度学习模型在复杂文本、视觉和多模态认知推理能力的数据集。该数据集包含2728个多选题,涉及26个不同的问题类别,主要来源于印度全国性的NTSE考试。NTSEBENCH旨在测试模型在不需要特定领域知识或死记硬背的情况下,解决问题的固有能力。数据集的创建过程包括从过往的NTSE试卷中提取问题,并通过OCR技术和人工校对进行数据清洗和处理。NTSEBENCH主要应用于评估和提升模型在认知推理任务中的表现,特别是在需要抽象和空间推理的视觉谜题解决方面。
NTSEBENCH is a dataset jointly created by the Indian Institute of Technology Guwahati, the University of Utah, and the University of Pennsylvania, designed to evaluate the capabilities of large deep learning models in complex textual, visual, and multimodal cognitive reasoning. It contains 2,728 multiple-choice questions spanning 26 distinct question categories, primarily sourced from India's national NTSE examination. NTSEBENCH aims to test a model's intrinsic problem-solving abilities without requiring specialized domain knowledge or rote memorization. The dataset construction process involves extracting questions from past NTSE examination papers, followed by data cleaning and processing using OCR technology and manual proofreading. NTSEBENCH is mainly applied to evaluate and enhance a model's performance on cognitive reasoning tasks, particularly in visual puzzle-solving that requires abstract and spatial reasoning.




