遇见数据集

ArchiveGene Corpus (v1.1): Uzbek Historical Archival Text Benchmark for NER, Coreference, and Genealogical Relation Extraction

收藏
Zenodo2026-06-12 更新2026-06-17 收录
官方服务:

资源简介:

ArchiveGene is a synthetic, multi-layer annotated corpus for information extraction from Uzbek historical archival texts. The corpus contains 1,000 documents split into training (700), validation (150), and test (150) sets. Each document is annotated across five layers: named entities, person mentions, coreference clusters, genealogical relation triples, and final structured tuples. The corpus is generated by a deterministic, template-based pipeline, enabling full reproducibility. It is designed as a benchmark for Named Entity Recognition, Coreference Resolution, and Genealogical Relation Extraction tasks. Format: JSON-L (one document per line). See README.md for the full schema and annotation guidelines.v1.1: Added UNCERTAIN mention type (8th type); extended document text.

提供机构:
Zenodo
创建时间:
2026-06-12
二维码
社区交流群
二维码
科研交流群
商业服务