hriaz/syntaxgym-hexatagged
收藏资源简介:
该数据集包含多个配置,每个配置针对不同的语言学或自然语言处理任务,可能用于评估模型在句法结构、语义合理性或心理语言学现象上的性能。数据集特征包括套件名称、项目编号、条件(包含条件名称、内容和区域信息)、预测结果以及多个以FORM、HEXATAG、T、NT为前缀的序列字段,这些字段可能代表不同模型或方法在合理或不合理条件下的输出。配置名称如center_embed、cleft、fgd等暗示任务涉及中心嵌入、裂句、填充语缺口等语言学结构。所有配置仅包含测试分割,用于模型评估。
This dataset includes multiple configurations, each targeting different linguistic or natural language processing tasks, likely for evaluating model performance on syntactic structures, semantic plausibility, or psycholinguistic phenomena. The features consist of suite name, item number, conditions (including condition name, content, and regions), predictions, and multiple sequence fields prefixed with FORM, HEXATAG, T, and NT, which may represent outputs from different models or methods under plausible or implausible conditions. Configuration names such as center_embed, cleft, and fgd suggest tasks involving linguistic structures like center embedding, cleft sentences, and filler-gap dependencies. All configurations contain only a test split for model evaluation.



