DWUG DE Sense: A data set of historical word sense annotations in German
收藏资源简介:
This data collection contains a subset of DWUG DE word usage data annotated with classical word sense definitions (DWUG DE Sense, see data/*/judgments_senses.csv). From these annotations aggregated and cleaned sense labels were derived (labels/*/labels_senses.csv). From these labels we derived additional binary semantic proximity labels between use pairs ('0' for different sense, '1' for same sense, labels/*/labels_proximity.csv) and change labels reflecting sense changes between the two time periods from which word usages were sampled (stats/*/stats_groupings.csv). The sense labels were derived from the sense annotation by removing instances where not at least 2/3 annotators agree on the label (maj_2/maj_3). Note that the binary proximity labels were derived from the sense annotation, and not directly judged by humans (in contrast to other WUG data sets). Note that consequently also the change scores EARLIER, LATER and COMPARE were not calculated directly from human judgments, but from the inferred binary proximity labels. Please find the code aggregating and cleaning the data, deriving proximity labels and deriving change labels in the WUG repository. Please find more information on the provided data in the paper referenced below. Version: 1.0.1, 01.11.2024. Correct or remove some normalization and lemmatization errors in the uses. Updated references. Reference Dominik Schlechtweg, Frank D. Zamora-Reina, Felipe Bravo-Marquez, Nikolay Arefyev. 2024. Sense Through Time: Diachronic Word Sense Annotations for Word Sense Induction and Lexical Semantic Change Detection. Language Resources and Evaluation. Dominik Schlechtweg. 2023. Human and Computational Measurement of Lexical Semantic Change. PhD thesis. University of Stuttgart.
本数据集包含DWUG DE词汇使用数据的子集,该子集附带经典词义定义标注(DWUG DE Sense,详见data/*/judgments_senses.csv)。基于上述标注,我们经聚合与清洗后得到了词义标注集(labels/*/labels_senses.csv)。基于上述词义标注集,我们进一步生成了词例对之间的二元语义邻近度标注(取值为0表示词义不同,取值为1表示词义相同,对应文件为labels/*/labels_proximity.csv),以及反映词例采样自两个时间段的词义变化标注(对应文件为stats/*/stats_groupings.csv)。 本次生成的词义标注集,是从原始词义标注中剔除未达到至少2/3标注者一致同意的条目后得到的(对应过滤规则为maj_2/maj_3)。请注意,本数据集的二元语义邻近度标注是基于词义标注推导而来,而非经人类直接标注(这与其他WUG数据集有所不同)。因此,EARLIER、LATER与COMPARE这三类变化分值也并非直接通过人工标注计算得到,而是基于推导得到的二元语义邻近度标注生成的。有关数据聚合清洗、邻近度标注生成以及变化标注生成的完整代码,请参见WUG代码仓库。 有关本数据集的更多详细信息,请参见下文列出的参考文献。 版本:1.0.1,发布日期:2024年11月1日。本次更新修正或移除了词例中的部分归一化与词形还原错误,并更新了参考文献列表。 参考文献 1. Dominik Schlechtweg、Frank D. Zamora-Reina、Felipe Bravo-Marquez、Nikolay Arefyev. 2024. Sense Through Time: Diachronic Word Sense Annotations for Word Sense Induction and Lexical Semantic Change Detection. Language Resources and Evaluation. 2. Dominik Schlechtweg. 2023. Human and Computational Measurement of Lexical Semantic Change. PhD thesis. University of Stuttgart.



