ENEIDE: A High Quality Silver Standard Dataset for Named Entity Recognition and Linking in Historical Italian
收藏资源简介:
ENEIDE (Extracting Named Entities from Italian Digital Editions) is a silver standard dataset for Named Entity Recognition and Linking (NERL) in historical Italian texts. The corpus comprises 2,111 documents semi-automatically extracted from two scholarly digital editions: Digital Zibaldone, the digital edition of Giacomo Leopardi's philosophical diary (1817-1832) and Aldo Moro Digitale, including the complete works of an Italian politician (1930-1978). ENEIDE contains over 8,000 entity annotations across multiple types (person, location, organization, literary work) linked to Wikidata identifiers. The dataset addresses the scarcity of annotated resources for historical Italian and provides a multi-domain, diachronic corpus suitable for evaluating NERL systems on humanistic and political texts with complex temporal and contextual features.



