A dataset of geographic entities and relationships from Song Dynasty texts on Lin'an
收藏资源简介:
Overview GEPR-LINAN is a manually annotated dataset for named entity recognition (NER) and relation extraction (RE) from Classical Chinese historical texts. The dataset focuses on geographical entities and spatial relationships described in texts about Lin'an (临安, present-day Hangzhou, Zhejiang Province, China), the capital of the Southern Song Dynasty (1127–1279 CE). Lin'an was one of the largest and most prosperous cities in the medieval world. Its spatial organisation — encompassing palaces, government offices, Buddhist and Taoist establishments, markets, gardens, waterways, and street networks — is documented in exceptional detail across a range of surviving local gazetteers and miscellaneous notes. These texts provide a foundation for computational reconstruction of the city's historical geography but pose significant challenges for NLP systems due to archaic vocabulary, elliptical syntax, and dense spatial expressions. The dataset provides annotations for 24 geographical entity types and 34 spatial and semantic relationship types, covering 4,920 sentences drawn from 18 text units drawn from 15 in-domain (IND) classical Chinese source works and one out-of-distribution (OOD) source. Metric Value Entity types 24 Relation types 34 Total sentences 4,920 Total entities 25,397 Total relations 17,881 Source texts (IND) 18 text units / 15 source works / 88 volumes Source texts (OOD) 1 work / 2 volumes IAA — Entity F1 86.38% IAA — Relation F1 77.82%
概述 GEPR-LINAN是一款面向古典中文历史文本的人工标注数据集,用于命名实体识别(Named Entity Recognition, NER)与关系抽取(Relation Extraction, RE)。该数据集聚焦于南宋(Southern Song Dynasty,1127–1279 CE)都城临安(Lin'an,今中国浙江省杭州市)相关文本中记载的地理实体与空间关系。 临安是中世纪世界规模最大且最为繁荣的城市之一。其空间布局涵盖宫殿、官署、佛道寺观、市集、园林、水道与街巷网络,现存的各类本地地方志与杂记中对其有着极为详尽的记载。此类文本可为该城市历史地理的计算重构提供坚实基础,但因文本存在古奥词汇、省略句法与密集空间表述等特点,给自然语言处理(Natural Language Processing, NLP)系统带来了显著挑战。 该数据集共标注了24种地理实体类型与34种空间及语义关系类型,涵盖来自15个域内(In-Domain, IND)古典中文源著作的18个文本单元,以及1个域外(Out-of-Distribution, OOD)源文本的总计4920个句子。 指标 数值 实体类型 24 关系类型 34 句子总数 4,920 实体总数 25,397 关系总数 17,881 域内源文本 18个文本单元 / 15部源著作 / 88卷 域外源文本 1部著作 / 2卷 标注者间一致性(实体)F1值 86.38% 标注者间一致性(关系)F1值 77.82%



