FOSSIL (Footnote-based Open-access SSH Scientific Instance Labels)
收藏资源简介:
FOSSIL数据集是一个面向法律与人文学科的多语言开放获取标注语料库,由葡萄牙ScienciaLAB和德国马克斯·普朗克法律史与法律理论研究所联合创建,旨在解决脚注中引文提取的难题。该数据集包含96篇学术文章,涵盖法律、人文、历史和社会科学等多个学科,涉及英语、葡萄牙语、意大利语和德语,时间跨度为1959年至2025年,共标注了约7,600条参考文献。其构建过程采用基于Grobid架构的PDF-TEI Editor工具,通过多人协作的五阶段标注工作流程完成,确保了数据的高质量和一致性。该数据集主要应用于法律与人文学科的引文解析、信息抽取和知识图谱构建,为训练和评估领域专用模型提供了关键资源,以应对脚注中混杂的参考文献、交叉引用和评论内容所带来的解析挑战。
FOSSIL Dataset is a multilingual open access annotated corpus tailored for law and humanities, co-developed by ScienciaLAB from Portugal and the Max Planck Institute for Legal History and Legal Theory in Germany. It aims to address the challenge of citation extraction from footnotes. The dataset comprises 96 academic articles spanning multiple disciplines including law, humanities, history and social sciences, covering four languages: English, Portuguese, Italian and German, with a time span from 1959 to 2025, and a total of approximately 7,600 annotated references. Its construction adopts the PDF-TEI Editor tool based on the Grobid architecture, and is completed through a five-stage collaborative annotation workflow involving multiple contributors, ensuring high data quality and consistency. This dataset is primarily applied to citation parsing, information extraction and knowledge graph construction in law and humanities, providing a critical resource for training and evaluating domain-specific models to tackle the parsing challenges posed by mixed references, cross-references and commentaries in footnotes.

- 1Digging Up Citations: FOSSIL, a Dataset and Workflow for Reference Extraction in Law and the Humanities葡萄牙ScienciaLAB; 德国马克斯·普朗克法律史与法律理论研究所 · 2026年




