British National Corpus
收藏资源简介:
该数据集是英国国家语料库,包含450对句子,用于研究语言模型是否能够捕捉到双宾语结构和介词宾语结构中受词的语义差异。数据集中的句子经过人工评估,确保在双宾语结构中受词被解释为人物,而在介词宾语结构中受词被解释为地点。这些数据被用于训练和评估语义特征库中的模型,该库可以投影上下文词嵌入到语义空间中,并分析语义解释的变化。
This dataset, derived from the British National Corpus, consists of 450 sentence pairs, and is developed to investigate whether language models can capture the semantic differences between the relevant objects in double-object constructions and prepositional-object constructions. All sentences in this dataset have been manually evaluated to ensure that the object in double-object constructions is interpreted as a person, while the object in prepositional-object constructions is interpreted as a location. These data are utilized to train and evaluate models within the semantic feature library, which can project contextual word embeddings into the semantic space and analyze variations in semantic interpretations.




