LVC_sentences_database
收藏资源简介:
该数据集是由庞培法布拉大学创建的用于探究语言模型中短语能力的最小对比对数据集,专注于英语轻动词结构与完整动词用法的区分。数据集包含47,600条精心构造的句子,涵盖make、take、give、have、receive等高频动词,通过控制句法框架和上下文生成轻动词与完整动词的对比序列。其构建过程基于CollFrEn语料库筛选候选轻动词组合,并人工验证确保名词的述谓性基础与动词的语义轻量性,同时匹配相应的完整动词-名词对以形成最小对比。该数据集主要应用于自然语言处理领域,旨在评估语言模型是否能够区分同一动词在轻动词结构与完整词汇用法之间的细微差异,为语言模型的短语能力探测提供标准化测试工具。
This minimal pair dataset, developed by Pompeu Fabra University for investigating the phrase-level competence of language models, focuses on distinguishing between English light verb constructions and full verb usage. Comprising 47,600 carefully constructed sentences covering high-frequency verbs including make, take, give, have, and receive, the dataset generates contrastive sequences of light verb and full verb structures by controlling syntactic frames and contextual settings. Its construction workflow screens candidate light verb combinations from the CollFrEn corpus, followed by manual validation to confirm the predicative foundation of the involved nouns and the semantic lightness of the target verbs, while matching corresponding full verb-noun pairs to form minimal contrast pairs. Primarily utilized in the field of natural language processing (NLP), this dataset aims to assess whether language models can discern subtle differences between the same verb when used in light verb constructions versus full lexical usage, serving as a standardized testbed for probing the phrase-level competence of language models.
数据集概述
- 数据集名称:LVC_sentences_database(轻动词结构句子数据库)
- 目的:收集包含轻动词结构(Light Verbs Constructions)的句子,供实验使用。
- 内容:提供可直接使用的句子数据集,以及用于生成数据集的脚本和材料。
数据获取
- 直接下载:访问 datasets 目录下载现成的数据集文件。
- 生成工具:在 LVC_sentences_generator 文件夹中,包含用于生成数据集的脚本和材料。




