Kangxi-Dictionary-V1.0
收藏资源简介:
康熙字典原子级结构化与多标准对齐数据集是一个针对经典古籍《康熙字典》进行深度清洗、多源标准对齐与数字化重构的高质量开源知识库。该数据集通过严格的脚本清洗、繁体规范化、笔画数纯数字转化以及异体字与地区标准对齐,将全量近47,000条字符数据转化为兼具学术考证价值和机器可读性的古汉字数字资产。数据规模为46,981条记录,涵盖字头、拼音、IDS字符结构树、地区标准说明(包括大陆GB、台湾CNS、日本JIS、韩国KS等)、所属集与部首、总画与余画(纯数字格式)、Unicode统一码、反切读音以及原典古释义。数据集采用JSON格式存储,每个条目包含唯一标识、字头、拼音、字结构、地区标准归属对照、字典收录归属与部首、笔画数、Unicode编码、传统反切音韵标注、康熙字典原释义等核心字段,并内嵌全局数字指纹和逐条明文系统版本标识以保护版权。数据集适用于文本分类、数字人文、自然语言处理等任务,特别支持古汉字研究、字符标准化对齐和历史文化分析。开源许可为CC BY-NC 4.0,仅限学术研究与非商业使用。
Atomically Structured and Multi-standard Aligned Kangxi Dictionary Dataset is a high-quality open-source knowledge base that conducts deep cleaning, multi-source standard alignment and digital reconstruction for the classic ancient Chinese book *Kangxi Dictionary*. This dataset converts nearly 47,000 full character entries into digital assets of ancient Chinese characters with both academic textual research value and machine readability, via rigorous script cleaning, traditional Chinese standardization, pure numeric conversion of stroke counts, and alignment between variant characters and regional standards. With a total of 46,981 records, the dataset covers character heads, pinyin, IDS character structure trees, regional standard descriptions (including mainland China GB, Taiwan CNS, Japan JIS, South Korea KS, etc.), affiliated collection and radical, total strokes and remaining strokes (pure numeric format), Unicode codes, fanqie pronunciation, and original definitions from the Kangxi Dictionary. Stored in JSON format, each entry in the dataset includes core fields such as unique identifier, character head, pinyin, character structure, regional standard attribution comparison, dictionary inclusion attribution and radical, stroke count, Unicode code, traditional fanqie phonetic annotation, original definitions from the Kangxi Dictionary, etc., and embeds a global digital fingerprint and per-entry plaintext system version identifier to protect copyright. The dataset is applicable to tasks such as text classification, digital humanities, natural language processing, and particularly supports ancient Chinese character research, character standardization alignment, and historical and cultural analysis. It is released under the CC BY-NC 4.0 open-source license, which is only allowed for academic research and non-commercial use.
数据集概述
项目名称:康熙字典原子级结构化与多标准对齐数据集 (KangXi Dictionary Structured Dataset)
核心定位:对《康熙字典》进行原子级深度清洗、多源标准对齐与数字化重构的高质量开源知识库,兼具学术考证价值与机器友好性。
数据规模与核心指标
- 总记录条目数:46,981 条
- 涵盖内容:字头、拼音、IDS 字符结构树、地区标准说明(大陆 GB、台湾 CNS、日本 JIS、韩国 KS 等)、所属集与部首、总画与余画(纯数字格式)、Unicode 统一码、反切读音、原典古释义。
- 版权保护:内嵌全局数字指纹(SHA-256)及逐条明文系统版本标识(
system_V1.0_ID)。
核心字段结构
数据集采用 JSON 格式存储,每个条目包含以下核心字段:
| 字段名 | 说明 |
|---|---|
id |
条目唯一检索标识 |
字 |
康熙字典原典字头 |
拼音 |
现代汉语拼音 |
字结构 |
Unicode 标准 IDS 原子结构树 |
字結構說明 |
地区标准归属对照 |
所屬集 / 所屬部 |
字典收录归属与部首 |
總畫 / 餘畫 |
经清洗的纯数字笔画数 |
統一碼 |
国际标准 Unicode 编码 |
讀音 |
传统反切与音韵标注 |
古釋義 |
原汁原味的《康熙字典》训诂释义 |
system_V1.0_ID |
逐条明文版权与系统版本指纹标识(位于条目底部) |
数字版权与开源声明
- 开源许可:采用 CC BY-NC 4.0 协议授权,仅限学术研究与非商业开源交流使用。
- 数字水印防护:内嵌全局加密数字指纹及逐条明文追踪 ID。
元数据信息
- 语言:中文(zh)
- 任务类型:文本分类(text-classification)
- 标签:kangxi、kangxi-dictionary、chinese-characters、digital-humanities、nlp、historical-text
- 数据规模:10K < n < 100K




