ArchiveGene Corpus (v1.1): Uzbek Historical Archival Text Benchmark for NER, Coreference, and Genealogical Relation Extraction
收藏资源简介:
ArchiveGene is a synthetic, multi-layer annotated corpus for information extraction from Uzbek historical archival texts. The corpus contains 1,000 documents split into training (700), validation (150), and test (150) sets. Each document is annotated across five layers: named entities, person mentions, coreference clusters, genealogical relation triples, and final structured tuples. The corpus is generated by a deterministic, template-based pipeline, enabling full reproducibility. It is designed as a benchmark for Named Entity Recognition, Coreference Resolution, and Genealogical Relation Extraction tasks. Format: JSON-L (one document per line). See README.md for the full schema and annotation guidelines.v1.1: Added UNCERTAIN mention type (8th type); extended document text.



