English Novel Compounds
收藏资源简介:
This release contains two datasets of novel English noun–noun compounds derived from an updated and cleaned version of the dataset associated with Learning to Predict Novel Noun-Noun Compounds (Dhar & van der Plas, 2019). The original study used temporally segmented data from the Google Books Ngram corpus to model and evaluate novel compound prediction across decades. Compared with the 2019 version, this updated release includes three main changes: (1) incorporation of Google Books v3, adding the 2010–2019 decade; (2) extraction restricted to noun sequences identified as compound relations by the spaCy dependency parser; and (3) removal of compounds that are part of a named entity. The release contains two tab-separated files: novel_compounds_2010.tsv, with compounds that are novel in the 2010s, and novel_compounds_2000.tsv, with compounds that were novel in the 2000s and are also attested in the 2010s. Each row reports the modifier, head, decade, and frequency count of the compound.



