GTC (Genocide Transcript Corpus)
收藏资源简介:
GTC(种族灭绝转录语料库)是由雷根斯堡大学创建的第一个种族灭绝相关法庭转录的标注语料库。该数据集包含1475条文本片段,来源于柬埔寨特别法庭(ECCC)、卢旺达国际刑事法庭(ICTR)和前南斯拉夫国际刑事法庭(ICTY)。数据集的创建旨在为社区提供一个参考语料库,建立新的分类任务基准,并探索领域内的迁移学习。GTC特别关注于标注那些描述暴力经历的证人陈述,这些陈述对于判断案件至关重要。数据集的应用领域主要集中在种族灭绝研究,旨在通过自动化工具减少人工研究的工作量,提高搜索效率。
The Genocide Transcript Corpus (GTC) is the first annotated corpus of genocide-related court transcripts developed by the University of Regensburg. This dataset contains 1475 text segments sourced from the Extraordinary Chambers in the Courts of Cambodia (ECCC), the International Criminal Tribunal for Rwanda (ICTR), and the International Criminal Tribunal for the former Yugoslavia (ICTY). The dataset was constructed to provide the research community with a reference corpus, establish benchmarks for novel classification tasks, and explore in-domain transfer learning. GTC specifically focuses on annotating witness statements that recount violent experiences, as such statements are critical for case adjudication. Its target application domains primarily center on genocide studies, aiming to reduce the workload of manual research and improve search efficiency through automated tools.



