遇见数据集

Chinese Characters Ontological Structure and Civilizational Coding Infrastructure (v1.0)

收藏
Zenodo2026-06-11 更新2026-06-12 收录
官方服务:

资源简介:

This dataset is the v1.0 release of the "Chinese Characters Ontological Structure and Civilizational Coding Infrastructure" — a full-system deep network knowledge infrastructure covering 9,795 Chinese characters from oracle bone script to modern characters, with Shuowen Jiezi (说文解字) as its core starting point. The database (SQLite format, 19 tables, ~186,000 rows) contains: - 9,795 Chinese characters with complete form-sound-meaning three-axis annotations - 43,137 typed cross-character relationships with 16 relationship type symbols and A/B/C evidence grading - 896 sound series (14,002 members), 1,730 form series (17,880 members), 48 meaning series (7,070 members) - 17,941 glosses from classical, Buddhist, and philological sources - 30,999 dialect pronunciations across 6 dialect points - 8,009 phonetic loan character (tongjiazi) relationships The dataset has produced original theoretical discoveries including: - Five paradigms of Chinese character encoding (Process 57.0%, Relation 37.6%, Structure, Origin, Fifth) - Eight typologies of sound-series conceptual genealogy - Eleven Laws of Chinese Character Encoding - Three practice-domain gene expression framework (Religious, Life, Production) Methodology: Characters are treated as multi-layered relational nodes. The seven-layer analysis framework spans from three-axis source tracing (Layer 1) through same-form/same-sound/same-structure comparison (Layers 2-5) to philosophical discovery (Layer 7). --- 本数据集是"中国汉字本体构造与文明编码底座"v1.0 版本——以《说文解字》为核心,覆盖从甲骨文到现代汉字的全系统深层网络知识基础设施。包含 9,795 个汉字、43,137 条类型化跨字 关系、三系网络(声符系/形符系/义符系)全覆盖。已支撑汉字编码五大范式、声系类型学、编码十一律等原创理论发现。

提供机构:
Zenodo
创建时间:
2026-06-11
二维码
社区交流群
二维码
科研交流群
商业服务