A Multiple-Choice Dataset for Evaluating Spanish Differential Object Marking in Large Language Models
收藏资源简介:
This repository contains two multiple-choice datasets designed to evaluate the linguistic competence of Large Language Models (LLMs) with respect to Differential Object Marking (DOM) in Spanish. DOM is a gradient, context-sensitive phenomenon whereby certain direct objects are preceded by the preposition a depending on semantic and referential properties of the object (animacy, definiteness) and the verb (telicity, affectedness, agentivity). Each dataset consists of 32 sentences (2⁵), balanced across five semantic categories coded as a binary feature matrix. Dataset 1 targets obligatory DOM contexts, presenting two options per item — one grammatically correct and one incorrect — to assess model accuracy. Dataset 2 targets optional DOM contexts, where both options are grammatically acceptable, allowing analysis of model preferences and sensitivity to the semantic factors that condition DOM variation in natural Spanish. All items were manually produced and reviewed by three expert linguists through multiple rounds of discussion.These resources are intended to support fine-grained linguistic evaluation of AI systems and reproducible research on Spanish grammar in NLP.



