遇见数据集

SemEval annotated dataset

收藏
SSH Open MarketPlace2023-10-27 更新2024-08-03 收录
官方服务:

资源简介:

This dataset contains the annotations for 40 Latin lemmas taken from the diachronic Latin corpus LatinISE. The dataset was originally created as Latin test data for SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection (Schlechtweg et al. 2020). The choice of the set of lexemes includes lexical units in which a change of meaning is observed in relation to Christianity and other socio-political changes in the late antiquity period. The annotation was performed manually: the annotators had to read each text snippet and assign the dictionary senses to each lemma in context on a graded scale, following a variation of the DuRel annotation framework (Schlechtweg et al. 2018): “1” indicated that the usage of the lemma in the text was unrelated to the dictionary sense, “2” indicated a distant relation between the two, “3” a close relation and “4” was used to indicate that the usage in the text completely overlapped with the dictionary sense. The label “0” was used when the annotator was not able to make a decision. References Schlechtweg, Dominik, im Walde Sabine Schulte & Stefanie Eckmann. 2018. Diachronic usage relatedness (DURel): A framework for the annotation of lexical semantic change. In Marilyn Walker, Heng Ji & Amanda Stent (eds.), The 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), 169–174. Stroudsburg, PA: Association for Computational Linguistics. Schlechtweg, Dominik, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky & Nina Tahmasebi. 2020. SemEval-2020 Task 1: Unsupervised lexical semantic change detection. In Aurelie Herbelot, Xiaodan Zhu, Alexis Palmer, Nathan Schneider, Jonathan May & Ekaterina Shutova (eds.), Proceedings of the fourteenth workshop on semantic evaluation, 1–23. Barcelona: International Committee for Computational Linguistics.

本数据集包含撷取自历时拉丁语语料库LatinISE的40个拉丁语词元(lemma)的标注数据。该数据集最初作为SemEval 2020任务1:无监督词汇语义变化检测(Schlechtweg等人,2020)的拉丁语测试数据集构建。本次选取的词项集合涵盖了古代晚期时期,伴随基督教及其他社会政治变革出现语义变化的词汇单位。 标注工作以人工方式完成:标注人员需阅读每一段文本片段,并参照经改进的DuRel标注框架(Schlechtweg等人,2018),依据分级评分量表为上下文语境中的每个词元分配词典义项。其中,等级“1”表示该词元在文本中的用法与词典义项无关联,“2”表示二者关联度较弱,“3”表示二者关联度较强,“4”表示该词元在文本中的用法与词典义项完全重合;当标注人员无法作出判定时,使用标签“0”。 参考文献 Schlechtweg, Dominik, Sabine Schulte im Walde 与 Stefanie Eckmann. 2018. 历时用法相关性(DURel):一种词汇语义变化标注框架. 载于Marilyn Walker、Heng Ji与Amanda Stent主编,《2018年北美计算语言学协会会议录:人类语言技术,第2卷(短文集)》,169–174页。宾夕法尼亚州斯特劳兹堡:计算语言学协会。 Schlechtweg, Dominik, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky与Nina Tahmasebi. 2020. SemEval-2020任务1:无监督词汇语义变化检测. 载于Aurelie Herbelot、Xiaodan Zhu、Alexis Palmer、Nathan Schneider、Jonathan May与Ekaterina Shutova主编,《第十四届语义评测研讨会论文集》,1–23页。巴塞罗那:国际计算语言学委员会。

创建时间:
2023-10-27
搜集汇总
数据集介绍
SemEval annotated dataset 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务