CENTERBENCH
收藏资源简介:
CENTERBENCH是一个包含9,720个理解问题的数据集,旨在测试语言模型是否真正理解句法结构或依赖于语义捷径。数据集包含360个中心嵌入句子,其中包含控制复杂性缩放和可能性/不可能性配对。每个句子都有六个理解问题,涉及表面理解、句法依赖和因果推理。数据集旨在帮助研究人员评估模型是否在处理复杂句子时放弃结构分析而转向语义捷径,从而提高模型的评估能力。
CENTERBENCH is a dataset consisting of 9,720 comprehension questions, designed to test whether language models truly understand syntactic structures or rely on semantic shortcuts. It contains 360 center-embedded sentences, with controlled complexity scaling and plausible/implausible pairs. Each sentence is paired with six comprehension questions covering surface-level comprehension, syntactic dependency, and causal reasoning. This dataset aims to help researchers evaluate whether models abandon structural analysis in favor of semantic shortcuts when processing complex sentences, thereby enhancing the rigor of model evaluation.

- 1The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for ShortcutsBrock University, St. Catharines, Canada & Emory University, Atlanta, USA · 2025年



