Compositionality Trend Prediction dataset
收藏资源简介:
本数据集由斯图加特大学研究团队构建,旨在探究德语和英语名词复合词的语义演变与组合性趋势。该数据集包含23个德语和26个英语目标复合词,通过历时语料库中每十年的上下文组合性评分,形成跨多个十年的组合性趋势数据。数据创建过程涉及从历时语料库中采样并人工标注组合性评分,以捕捉复合词含义的渐变过程。该数据集主要应用于自然语言处理领域,特别是词汇语义变化检测任务,旨在解决复合词组合性随时间变化的量化分析问题,为语义演变建模提供细粒度的时间序列评估基准。
This dataset was constructed by a research team from the University of Stuttgart, aiming to explore the semantic evolution and compositionality trends of German and English noun compounds. It includes 23 target German noun compounds and 26 target English noun compounds, forming compositional trend data spanning multiple decades based on decade-by-decade contextual compositionality scores extracted from the diachronic corpus. The data creation process involves sampling from the diachronic corpus and manually annotating compositionality scores to capture the gradual semantic shift of compound words. This dataset is primarily applied in the field of Natural Language Processing (NLP), particularly for lexical semantic change detection tasks. It aims to address the quantitative analysis of temporal changes in the compositionality of noun compounds, and provides fine-grained time-series evaluation benchmarks for semantic evolution modeling.

- 1Losing My Composure: Predicting Compositionality Over Time斯图加特大学 · 2026年



