NtoN Construction Dataset
收藏资源简介:
该数据集是针对英语中的NtoN结构(即名词+介词+名词)构建的,由乔治城大学的研究者创建。数据集包含了从COCA中提取的6599个实例,涵盖了两种语义类型:连续性和并置性。数据集在构建过程中,研究者通过固定窗口提取、分词、排除干扰项等步骤,最终形成了经过人工标注的、用于研究BERT模型对NtoN结构理解能力的数据集。该数据集旨在解决自然语言处理中对特定语言结构理解的问题。
This dataset was constructed for the N-to-N structure in English (i.e., Noun + Preposition + Noun) and was developed by researchers from Georgetown University. It contains 6599 instances extracted from COCA, covering two semantic types: continuity and juxtaposition. During the construction process, researchers adopted steps including fixed-window extraction, tokenization, and interference item elimination, and finally formed a manually annotated dataset for investigating BERT's ability to understand the N-to-N structure. This dataset aims to address the issue of understanding specific linguistic structures in natural language processing.

- 1Construction Identification and Disambiguation Using BERT: A Case Study of NPN乔治城大学 · 2025年



