遇见数据集

TCMNSCLC: A Real-world Dataset for Chinese Medicine Reasoning on Non-small-cell Lung Cancer

收藏
Zenodo2026-06-29 更新2026-08-13 收录
官方服务:

资源简介:

TCM4NSCLC is a real-world dataset developed for traditional Chinese medicine (TCM) reasoning in non-small-cell lung cancer (NSCLC). The dataset consists of structured clinical cases curated from real-world medical records and annotated by experienced TCM experts. Each case contains comprehensive patient information, including clinical characteristics, TCM syndrome differentiation, treatment principles, herbal decoction prescriptions, and Chinese patent medicine recommendations. The dataset is designed to facilitate research on large language models (LLMs) for TCM clinical reasoning, clinical decision support, prescription generation, and explainable medical artificial intelligence. The dataset provides the following information for each clinical case: Structured clinical case descriptions TCM syndrome differentiation labels Treatment principles Herbal decoction prescriptions Chinese patent medicine recommendations The dataset contains the following files: File Description TCM4NSCLC_Original_Clinical_Dataset.xlsx Original structured clinical dataset collected from real-world medical records. Compared with the JSON files, this file additionally includes patient demographic and visit information, such as patient_id, sex, age, accrual_type, and visit_date. It serves as the source dataset before splitting into training, validation, and test sets. train_num_3032.json Training set containing 3,032 clinical cases for model training. valid_num_379.json Validation set containing 379 clinical cases for hyperparameter tuning and model selection. test_num_379.json Test set containing 379 clinical cases for final evaluation. example_en.json An English example demonstrating the dataset format and field definitions. schema.json JSON schema describing the structure and data types of each field. README.md Documentation describing the dataset and usage instructions. The released dataset is divided into three subsets: Split Number of Cases Train 3,032 Validation 379 Test 379 Total 3,790 The training, validation, and test sets are generated from the original structured clinical dataset. The original Excel file additionally preserves patient demographic and visit-related metadata that are not included in the released JSON files for model training.

提供机构:
Zenodo
创建时间:
2026-06-29
二维码
社区交流群
二维码
科研交流群
商业服务