Structured Dataset of Traditional Efficacies for Chinese Herbal Medicines
收藏资源简介:
Description: This dataset provides a structured and computable resource of traditional therapeutic effects (efficacies) for Chinese herbal medicines. Its main purpose is to transform classical TCM efficacy descriptions from natural-language expressions into standardized herb-efficacy relations that can be directly used for knowledge graph construction, formula efficacy prediction, herb-symptom association mining, and AI-based reasoning over traditional therapeutic knowledge. The dataset links Chinese herbs to explicit and rule-expanded efficacy nodes. Explicit efficacy nodes are directly structured from authoritative TCM textual sources, while expanded efficacy nodes are inferred through rule-based semantic reasoning over a TCM knowledge graph. In this way, the dataset not only records what therapeutic effects are explicitly stated for each herb, but also provides computable extensions of those effects to related symptoms, patterns, sub-patterns, and pathogenesis-related efficacy concepts. Version 2.0 Release Notice: This is Version 2.0 of the Structured Dataset of Traditional Efficacies for Chinese Herbal Medicines. This release reorganizes the previous single-table dataset into separate herb-efficacy relation files and rule-specific evidence files, improving transparency, traceability, and downstream computability. A new field, Efficacy_Node_Type, is included to distinguish explicit efficacy nodes from rule-based expanded nodes generated by Rule A, Rule B, Rule C, and Rule D. Upstream Knowledge Graph Dependency Note: The rule-based expanded efficacy relations in this dataset were generated based on the updated Version 2.0 of the related TCM knowledge graph. Because the upstream knowledge graph was revised and further standardized, the semantic reasoning results produced by Rules A-D were also recalculated and updated accordingly. Therefore, the relation counts and expanded efficacy results in this version may differ from those in the previous release. The related knowledge graph is available at: https://doi.org/10.5281/zenodo.20061549 Dataset Content: The dataset consists of five herb-efficacy relation files and four rule-specific evidence files. Herb-Efficacy Relation Files: 1. explicit_herb_efficacy_relations_v2.csv: Directly curated herb-to-efficacy relations for explicit efficacy nodes. 2. expanded_herb_efficacy_relations_ruleA_v2.csv: Herb-to-efficacy relations expanded through Rule A. 3. expanded_herb_efficacy_relations_ruleB_v2.csv: Herb-to-efficacy relations expanded through Rule B. 4. expanded_herb_efficacy_relations_ruleC_v2.csv: Herb-to-efficacy relations expanded through Rule C. 5. expanded_herb_efficacy_relations_ruleD_v2.csv: Herb-to-efficacy relations expanded through Rule D. Evidence Files: 1. herb_efficacy_expansion_evidence_ruleA_v2.csv: Rule A evidence table documenting the reasoning paths for Rule A expansions. 2. herb_efficacy_expansion_evidence_ruleB_v2.csv: Rule B evidence table documenting the reasoning paths for Rule B expansions. 3. herb_efficacy_expansion_evidence_ruleC_v2.csv: Rule C evidence table documenting the reasoning paths for Rule C expansions. 4. herb_efficacy_expansion_evidence_ruleD_v2.csv: Rule D evidence table documenting the reasoning paths for Rule D expansions. Key Statistics: The five herb-efficacy relation files contain 34,133 relation records in total, covering 483 unique herbs and 2,564 unique target efficacy nodes. After deduplication by source_id and target_id, the dataset contains 33,859 unique herb-efficacy pairs. The four rule-specific evidence files contain 27,876 evidence records in total. Data Structure: The relation files are provided in CSV format with UTF-8 encoding. The shared fields include source_id, source_ENGLISHNAME, source_LABEL, target_id, target_ENGLISHNAME, target_LABEL, TYPE, and Efficacy_Node_Type. The TYPE field represents the directed has_effect relation between herbs and efficacy nodes. The Efficacy_Node_Type field indicates whether the entry belongs to explicit efficacy nodes or rule-based expanded nodes. The evidence files provide rule-specific provenance information for expanded efficacy relations. Because Rules A-D involve different semantic reasoning paths, the evidence files contain different intermediate-node fields. The number of evidence records may differ from the number of final expanded herb-efficacy relations because a single herb-efficacy relation may be supported by multiple reasoning paths or intermediate evidence nodes. Potential Applications: This resource is intended for researchers in TCM informatics, biomedical knowledge graphs, and AI-based prescription analysis. It supports formula efficacy prediction, herb-efficacy relation mining, symptom-herb association analysis, knowledge graph enrichment, semantic reasoning, and computational analysis of traditional therapeutic knowledge. Related Publication: Yuanbai L, Fangzhou L, Yihao L, Yu D, Meng L, Qin Q, Yang Y, Hongming M. A Knowledge Graph-Driven Hypergeometric Efficacy Prediction Model for Classical Traditional Chinese Herbal Formulas. Methods Inf Med. 2026 Apr 7. doi: 10.1055/a-2841-4549. Epub ahead of print. PMID: 41895302. Contact:For questions, please contact:LI Yuanbai: liyuanbai126@126.com



