遇见数据集

Labeled entities in judgements for Legal Named Entity Recognition training in Spanish

收藏
Zenodo2025-12-15 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains 125 labeled texts from judgments delivered by the Spanish Supreme Court of Justice. Each text is labelled for a Named Entity Recognition model training. The dataset does not contain original texts as they cannot be redistributed without permission. However, all the texts can be retrieved from the official source, the CENDOJ platform of the Spanish General Council of the Judiciary (https://www.poderjudicial.es/search/indexAN.jsp). All the elements of the dataset contain a ROJ id (following the European Case Law Identifier format). Each element of the dataset follows this structure: { "id": Id of the document. "label": List of entities. Each entity is a list with three elements [Start character, end character, label]. "roj": Id of the judgment in the CENDOJ repository. "offset": Character where the background of the judgment starts. } The offset marks the beggining of the background in the document, which is the part of the text that has been labelled. There are 4 different entities labelled: Entity Tag Description Law LAW References to Spanish laws and decrees. European Directive EU_DIR References to European Directives. Article ART References to articles of said laws and directives. Procedure identifier PROC_ID References to other non legislative documents (other judgments, edicts, etc.). Globally there are 1,620 LAW entities, 67 EU_DIR entities, 1,343 ART entities and 885 PROC_ID entities.

提供机构:
Zenodo
创建时间:
2025-12-15
二维码
社区交流群
二维码
科研交流群
商业服务