遇见数据集

Manual Topic Annotation of German Novels and Parlament Protocols by multiple Annotators

收藏
Zenodo2021-07-07 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This dataset was created in the research project hermA and contains topic annotations for 960 sentences, half of which were taken from transcripts of the German Bundestag and half from recent German novels. For each sentence, 30 different annotators evaluated whether illness is addressed, how central the topic is, if so, and how certain they are in the annotation. The dataset contains the following columns: <strong>item_id </strong>(for each sentence) <strong>worker_id </strong>(for each annotator) <strong>worker_group </strong>(either "crowdworker" or "student") <strong>corpus </strong>(either "protocol_corpus"=transcript corpus or "novel_corpus"=fiction corpus) <strong>text_source</strong> <strong>previous_sentences </strong>(in the text_source) <strong>target_sentence</strong> <strong>following_sentences </strong>(in the text_source) <strong>semantic_field_token </strong>(if existing) <strong>semantic_field_status </strong>(either "True" or "False") <strong>1_wird_im_fett_gedruckten_satz_krankheit_thematisiert </strong>(topic annotations: either "ja" or "nein") <strong>1b_wie_zentral_ist_das_thema_krankheit_im_fettgedruckten_satz </strong>(topic centrality: "NaN", "krankheit_kommt_eher_am_rande_des_satzes_vor" or "krankheit_ist_das_zentrale_thema_des_satzes") <strong>2_wie_sicher_bist_du_dir_bei_der_antwort_zu_frage_1_ </strong>(annotation certainty: "sehr_sicher", "eher_sicher", "eher_unsicher" or "sehr_unsicher") We use the annotations to model ambiguity in: Andresen, Melanie; Vauth, Michael &amp; Zinsmeister, Heike. 2020. Modeling Ambiguity with Many Annotators and Self-Assessments of Annotator Certainty. <em>Proceedings of 14th Linguistic Annotation Workshop</em>.

本数据集诞生于hermA研究项目,共包含960个句子的主题标注数据。其中半数句子取自德国联邦议院会议记录文稿,另一半取自当代德国小说作品。 针对每一句子,共有30名不同标注者完成三项评估:是否涉及疾病主题、该主题在句中的核心程度,以及自身标注的置信度。 数据集包含以下字段: - `item_id`:单句唯一标识符 - `worker_id`:标注者编号 - `worker_group`:标注者组别,分为「众包标注者(crowdworker)」与「学生(student)」两类 - `corpus`:语料库(corpus)来源,可选值为`protocol_corpus`(会议记录语料库)与`novel_corpus`(小说语料库,即虚构文本语料库) - `text_source`:文本来源 - `previous_sentences`:目标句在其文本来源中的上下文前文 - `target_sentence`:目标句子 - `following_sentences`:目标句在其文本来源中的上下文后文 - `semantic_field_token`:语义场标记(semantic_field_token,若存在) - `semantic_field_status`:语义场标记状态,取值为"True"或"False" - `1_wird_im_fett_gedruckten_satz_krankheit_thematisiert`:主题标注项,取值为"ja"(是)或"nein"(否),用于标注粗体句是否涉及疾病主题 - `1b_wie_zentral_ist_das_thema_krankheit_im_fettgedruckten_satz`:主题核心度标注项,取值为"NaN"、`krankheit_kommt_eher_am_rande_des_satzes_vor`(疾病仅作为句中次要主题提及)或`krankheit_ist_das_zentrale_thema_des_satzes`(疾病为句子核心主题) - `2_wie_sicher_bist_du_dir_bei_der_antwort_zu_frage_1_`:标注置信度项,取值为"sehr_sicher"(非常确信)、"eher_sicher"(较确信)、"eher_unsicher"(较不确信)或"sehr_unsicher"(完全不确信) 本数据集的标注结果被用于以下研究中的歧义建模工作:Andresen, Melanie; Vauth, Michael & Zinsmeister, Heike. 2020. 《多标注者与标注者置信度自评视角下的歧义建模》,收录于*第14届语言标注工作坊论文集(Proceedings of 14th Linguistic Annotation Workshop)*。

提供机构:
Zenodo
创建时间:
2020-10-14
二维码
社区交流群
二维码
科研交流群
商业服务