ANNOTARES
收藏资源简介:
ANNOTARES是一个专为提取德国法律文本逻辑结构而设计的数据集,由卡尔斯鲁厄应用科学大学创建。该数据集包含539条句子,超过21,500个标注的token,覆盖三部德国联邦法律(BDSG、BAföG、BauGB),并以跨度级别标注了法律条件(Tatbestand)和法律后果(Rechtsfolge)。数据集的创建过程由六名标注员协作完成,采用重叠标注策略并辅以“不确定”标记以排除歧义样本,最终通过自动化工具与人工校正确保高标注质量。该数据集旨在推动法律文本的自动化逻辑结构分析,为法律知识图谱、检索增强生成及神经符号推理等下游任务奠定基础,从而解决现有方法难以精确解析法律规范中条件与后果关系的核心挑战。
ANNOTARES is a dataset specifically designed for extracting the logical structure of German legal texts, created by Karlsruhe University of Applied Sciences. This dataset contains 539 sentences and over 21,500 annotated tokens, covering three German federal laws (BDSG, BAföG, BauGB), and annotates legal conditions (Tatbestand) and legal consequences (Rechtsfolge) at the span level. The dataset's development was a collaborative effort by six annotators, who adopted an overlapping annotation strategy supplemented with "uncertain" markers to exclude ambiguous samples, and ultimately ensured high annotation quality through automated tools and manual verification. This dataset aims to promote automated logical structure analysis of legal texts, lay a foundation for downstream tasks such as legal knowledge graphs, retrieval-augmented generation, and neural-symbolic reasoning, thereby addressing the core challenge that existing methods struggle to accurately parse the relationship between conditions and consequences in legal norms.
- 1ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts卡尔斯鲁厄应用科学大学 · 2026年



