遇见数据集

Mingzheng - Reproducibility Data Package (v3)

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

Mingzheng - Reproducibility Data Package (v3) This research resource provides selected model code, derived research materials and aggregate evaluation outputs from the Mingzheng multimodal syndrome-element classification project. The 15 unchanged frozen model checkpoints are hosted separately at https://doi.org/10.5281/zenodo.20111083 in mingzheng_checkpoints.zip and are not duplicated in this update. Mingzheng integrates dual-view tongue images, clinician-recorded tongue and pulse findings, inquiry text and language-model-derived evidence for conditional multi-label classification of phlegm-dampness, blood stasis and yin deficiency. The intended-use population has cancer with sleep disturbance and a clinician-confirmed background pattern of liver-qi stagnation with spleen deficiency. Development included 478 participants from five clinical sites; prospective temporal and independent-hospital intended-use validation included 66 Hangzhou and 73 Ningbo participants. A paired reader study included 19 trainees and 47 target-positive cases. The lightweight archive contains pseudonymous development split metadata, text-stripped LLM confidence scores and parser status, aggregate development and prospective-validation results, reader-by-case bootstrap and missingness summaries, clinical lexical/parser checks, a minimal model runtime, selected analysis scripts and refreshed public TCM-SD exploratory-adapter artifacts. It also provides the exact external weight-file link, size, MD5 and SHA-256 checksums, a verification utility and instructions for combining the materials. The published Zenodo weight-file size and MD5 were matched to the local frozen archive on 29 September 2026. No model retraining or numerical-result changes were introduced. Download Mingzheng_Reproducibility_v3_Light_20260929.zip and extract it normally. For model inference, also download mingzheng_checkpoints.zip from the version-specific DOI above and verify it with the included verify_weights.py utility. The refreshed TCM-SD materials are included in the lightweight archive; the earlier TCM-SD archive does not need to be downloaded separately. No raw hospital clinical records, tongue images, pulse waveforms, inquiry transcripts, clinical free-text LLM responses, patient-level Ningbo predictions or reader response matrices are included. Pseudonymous participant-level development splits and numeric LLM scores are included. The TCM-SD component contains public benchmark-derived narrative text and generated output, not restricted hospital records. The complete training and analysis repository is maintained privately. Requests for controlled access to additional code or suitably de-identified clinical derivatives can be directed to the contact on the hosting record; clinical data access remains subject to institutional approval and applicable data-sharing agreements. File-level terms apply: original research materials are CC BY 4.0; software is MIT; public TCM-SD-derived data and documentation are CC BY-NC-SA 4.0. See LICENSES.md for attribution and scope. The materials support inspection and reuse of the linked frozen model but do not reproduce patient-level hospital analyses without access to restricted inputs. Cite this record for the updated materials and https://doi.org/10.5281/zenodo.20111083 when using the model weights. The earlier public TCM-SD extension is https://doi.org/10.5281/zenodo.20382160.

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务