遇见数据集

De-identified dataset and Python analysis code: genotype-derived drug-response phenotypes in 908 Chinese patients with type 2 diabetes

收藏
Zenodo2026-07-28 更新2026-08-01 收录
官方服务:

资源简介:

This record contains the de-identified individual-level dataset and the complete Python analysis code for the study "Genotype-derived drug-response phenotypes in 908 Chinese patients with type 2 diabetes: a retrospective single-centre study". CONTENTS. Database_deidentified.csv is the de-identified dataset covering 908 patients: age, sex, genotypes at the seven analysed loci, and the pharmacogenomic phenotypes as issued by the testing laboratory. DATA_DICTIONARY.md gives a column-by-column description of that dataset. The scripts folder holds six Python scripts that reproduce every table and figure in the paper, and requirements.txt lists the Python dependencies. PRIVACY. Patient names are not included. Laboratory accession numbers have been replaced with sequential study codes (P0001-P0908); the linking key is held by the study team and is not published. The study was approved by the Medical Ethics Committee of Zaozhuang Mining Group Central Hospital (approval number 2022-伦理审查-02, granted 27 April 2022), which waived the requirement for individual informed consent for this retrospective analysis of de-identified records. NOTES ON THE DATA. The phenotype columns hold the interpretation as issued in the original clinical report, not a value recomputed from the genotype. For 23 of the 908 records this differs from the standardized rule published as Table 1 of the paper; those records are listed in S7 Table of the paper and were deliberately left as issued. Separately, the cohort was tested on two successive versions of the genotyping panel, a four-locus panel (427 patients) and a seven-locus panel (481 patients), so missingness at rs4149056, rs7756992 and rs2066865 is structural rather than random. The two panels also applied different interpretive rules to predicted sulfonylurea efficacy. ANALYSES. Genotype-to-phenotype mapping for seven SNPs across six genes (TCF7L2, IRS1, CDKAL1, CYP2C9, SLCO1B1, FGG); eight predicted clinical phenotypes; a weighted polygenic risk score built from fixed literature-based locus weights; and gene-gene interaction analysis with likelihood ratio tests. REQUIREMENTS. Python 3.9 or later with pandas, numpy, scipy, statsmodels, scikit-learn, matplotlib, seaborn and plotnine. CHANGES FROM VERSION 1.0.0. This version adds the de-identified dataset and its data dictionary, where version 1.0.0 contained code only. It corrects the description of the polygenic risk score, which uses fixed literature-based locus weights rather than study-derived coefficients. The polygenic risk score script no longer copies identifying columns of the input file into its per-patient output. The archive is distributed as .zip rather than .rar so that it opens without proprietary software.

提供机构:
Zenodo
创建时间:
2025-12-15
二维码
社区交流群
二维码
科研交流群
商业服务