LaCour!
收藏资源简介:
LaCour!是欧洲人权法院文本口头论据的第一个语料库,包含154个完整听证会(210万字符,超过267小时视频素材),涵盖英语、法语及其他法庭语言,每个听证会均与相应的最终判决文件关联。除了从视频转录并部分手动校正的文本外,还提供句子级时间戳和手动标注的角色及语言标签。此外,还展示了LaCour!在一系列初步实验中探索问题与异议意见之间互动的用途。除了在法律NLP中的应用,希望法律学生或其他感兴趣的团体也能将LaCour!作为学习资源,该数据集在https://huggingface.co/datasets/TrustHLT/LaCour上自由提供多种格式。
LaCour! is the first corpus of oral arguments from the European Court of Human Rights, comprising 154 complete hearings (2.1 million characters, over 267 hours of video material), covering English, French, and other courtroom languages. Each hearing is associated with its corresponding final judgment document. In addition to the text transcribed from video and partially manually corrected, sentence-level timestamps and manually annotated roles and language tags are provided. Furthermore, LaCour! demonstrates its utility in a series of preliminary experiments exploring the interaction between questions and dissenting opinions. Beyond its application in legal NLP, it is hoped that law students or other interested groups can use LaCour! as a learning resource. The dataset is freely available in various formats at https://huggingface.co/datasets/TrustHLT/LaCour.
数据集概述
数据集名称
- LaCour! Corpus
数据集内容
- LaCour! Corpus 包含154个欧洲人权法院听证会的完整文本记录,总计2.1百万字符,来自超过267小时的视频资料,涵盖英语、法语及其他法院语言。
数据集结构
-
子集:transcripts
- 包含154个听证会的文本记录。
- 提供两种格式:
.xml和.txt。 - 两种格式均包含以下信息:
- webcast_id
- Role
- Name
- Begin
- End
- Language
- text
-
子集:documents
- 包含与听证会相关的所有文档信息,这些文档来自HUDOC数据库,通过应用号与听证会关联。
- 每个文档实例包含以下信息:
- id
- webcast_id
- hearing_date
- hearing_title
- hearing_type
- appno
- case_id
- case_name
- case_url
- type
- typedescription
- document_date
- collection
- importance
- court
- issue
- represented_by
- respondent
- articles
- strasbourg_caselaw
- external_sources
- conclusion
- separate_opinion
- judges
- ecli
数据集用途
- 用于研究欧洲人权法院中的论证,特别是听证会中的问题与异议意见之间的相互作用。
- 作为法律自然语言处理的研究资源。
- 作为法律学生或其他感兴趣方的学习资源。
数据集访问
- 数据集可在以下链接免费访问:Huggingface Dataset
联系人
- Lena Held
- 邮箱:lena.held@tu-darmstadt.de




