Tao-Shanghan-Corpus: A Structured Corpus of Historical Shanghan Literature with Clauses, Commentaries and Textual Variants
收藏资源简介:
The Historical Shanghan Corpus (Tao-Shanghan-Corpus) v1.0 is a provenance-aware, machine-readable dataset of the historical *Shanghan Lun* (伤寒论) tradition. The corpus represents canonical clauses, historical commentaries, textual variants, and inter-textual relations as separate, addressable data layers linked to catalogued historical works and source-level provenance. The release contains 681 clause records, including 398 canonical clauses, 2,958 aligned historical commentaries, 616 textual variants, 4,286 inter-textual relations, 4,255 graph nodes, and 4,255 unified text records. The upstream corpus comprises 57 catalogued historical works represented by 425 source text files (4,763,991 characters). Structured records are distributed in CSV and JSONL with a field-level data dictionary, source provenance, file checksums, and computational validation reports. Computational validation reports are included with the release. The release includes the 425 upstream source transcription files. The underlying works are pre-modern. The current jicheng.tw official copyright statement checked on 2026-08-11 releases its editorial contributions to public-domain texts under CC0, subject to any page-specific notice. Structured dataset layers are CC BY 4.0. This dataset is intended for historical-text, philological, digital-humanities, corpus-linguistics, and knowledge-representation research and is not intended for clinical guidance.



