English Novel Compounds
收藏资源简介:
This release contains two datasets of novel English noun–noun compounds derived from an updated and cleaned version of the dataset associated with Learning to Predict Novel Noun-Noun Compounds (Dhar & van der Plas, 2019). The original study used temporally segmented data from the Google Books Ngram corpus to model and evaluate novel compound prediction across decades. Compared with the 2019 version, this updated release includes three main changes: (1) incorporation of Google Books v3, adding the 2010–2019 decade; (2) extraction restricted to noun sequences identified as compound relations by the spaCy dependency parser; and (3) removal of compounds that are part of a named entity. The release contains two tab-separated files: novel_compounds_2010.tsv, with compounds that are novel in the 2010s, and novel_compounds_2000.tsv, with compounds that were novel in the 2000s and are also attested in the 2010s. Each row reports the modifier, head, decade, and frequency count of the compound.
本发布版本包含两份英语新颖名词-名词复合词数据集,其源自《学习预测新颖名词-名词复合词》(Learning to Predict Novel Noun-Noun Compounds,Dhar与van der Plas,2019)关联数据集的更新与清洗版本。原始研究采用谷歌图书Ngram语料库(Google Books Ngram corpus)的时序分段数据,对跨年代的新颖复合词预测任务进行建模与评估。 相较于2019年版本,本次更新发布主要包含三项改进:(1)纳入谷歌图书v3(Google Books v3)数据集,新增2010-2019年代的数据;(2)提取范围限定为经spaCy依存句法分析器标记为复合词关系的名词序列;(3)移除属于命名实体的复合词。 本发布版本包含两份制表符分隔文件:分别为记录2010年代新颖复合词的novel_compounds_2010.tsv,以及记录2000年代首次出现且在2010年代仍有佐证的新颖复合词的novel_compounds_2000.tsv。每行数据分别记录复合词的修饰语、中心语、所属年代以及出现频次。



