xpertsystems/hc-gen-008-sample
收藏资源简介:
HC-GEN-008是一个合成基因组变异数据集样本,包含变异、基因型、多基因风险评分(PRS)表型、药物基因组学和转录组学数据。该数据集是完整HC-GEN-008产品(企业默认:200,000个样本×2,000,000个变异)的缩小版,具体为500个样本×5,000个变异。数据完全合成,由基准优先的模拟引擎生成,不包含任何真实的基因组、个体或临床数据,基因名称和变异ID均为虚构。数据集涵盖5个祖先群体(EUR、AFR、EAS、SAS、AMR),包括七个关系型CSV表格:变异注册表(含染色体位置、等位基因频率、基因功能影响等)、群体元数据、基因型矩阵(符合Hardy-Weinberg平衡)、表型结果(如疾病案例对照、BMI等性状)、药物基因组学(代谢器类别和毒性标志)、转录组学(基因表达和eQTL链接)以及队列摘要。数据集校准锚点基于真实基因组数据库(如1000 Genomes/gnomAD),确保变异类型、等位基因频率、罕见变异率、致病性率等指标在目标范围内。用途包括GWAS、多基因风险评分建模、群体遗传学分析、药物基因组学响应建模、eQTL整合和变异注释基准测试。但存在限制:如HWE计算硬编码、PRS R²被高估、GWAS信号为代理、规模缩减、合成位点虚构,以及仅边际校准而非完整联合保真度。数据集仅用于机器学习开发、基准测试、模式原型设计和教育,严禁用于临床或患者护理。
HC-GEN-008 is a synthetic genomic variation dataset sample containing variant, genotype, polygenic risk score (PRS) phenotype, pharmacogenomics, and transcriptomics data. This dataset is a scaled-down version of the full HC-GEN-008 product (default enterprise setting: 200,000 samples × 2,000,000 variants), specifically consisting of 500 samples × 5,000 variants. The data is fully synthetic, generated by a benchmark-prioritized simulation engine, and does not contain any real genomic, individual, or clinical data; all gene names and variant IDs are fictitious. The dataset covers 5 ancestry populations (EUR, AFR, EAS, SAS, AMR), and includes seven relational CSV tables: Variant Registry (including chromosomal positions, allele frequencies, gene functional impacts, etc.), Population Metadata, Genotype Matrix (compliant with Hardy-Weinberg Equilibrium), Phenotype Outcomes (e.g., disease case-control status, BMI and other traits), Pharmacogenomics (metabolizer categories and toxicity markers), Transcriptomics (gene expression and eQTL links), and Cohort Summary. The dataset’s calibration anchors are based on real genomic databases such as 1000 Genomes/gnomAD, ensuring that metrics including variant types, allele frequencies, rare variant rates, and pathogenicity rates fall within target ranges. Its applications include GWAS, polygenic risk score modeling, population genetics analysis, pharmacogenomics response modeling, eQTL integration, and variant annotation benchmarking. However, there are limitations: hard-coded HWE calculations, overestimated PRS R², proxy GWAS signals, reduced scale, fictitious synthetic loci, and only marginal calibration rather than full joint fidelity. This dataset is solely intended for machine learning development, benchmarking, model prototyping, and education, and is strictly prohibited for use in clinical or patient care.




