Symphonym: Universal Phonetic Embeddings for Cross-Script Name Matching - Models and Evaluation Data
收藏资源简介:
This dataset contains the complete trained models, vocabularies, evaluation results, and supporting materials for Symphonym v7, a neural system for cross-script name matching using teacher-student distillation.Key Features: Trained models (Phase 1–3) totalling ~300 MB 113,280-character vocabulary covering 20 major writing systems 66.9M toponym training corpus from GeoNames, Wikidata, and Getty TGN MEHDIE benchmark evaluation results (Recall@K, MRR across 5 testsets) 11,723 cross-script pair test results across 170+ script combinations 102 custom Epitran G2P extensions for 102 language-script pairs Models Included: phase1_best.pt (12 MB) - Teacher network trained on PanPhon phonetic features phase2_best.pt (96 MB) - Student network after contrastive distillation phase3_best.pt (96 MB) - Final model with hard negative mining final_model.pt (96 MB) - Production checkpoint Evaluation Results: - MEHDIE benchmark: Medieval Arabic-Latin toponym matching (5 testsets)- Cross-script pairs: 13,660 validated pairs across 1,366 script combinations- Metrics: Recall@1/5/10/20, MRR, precision, F-scoresUse Cases: - Historical gazetteer reconciliation- Multilingual named entity linking- Cross-script information retrieval- Phonetic similarity search- Historical name variant identification



