Decoding sequence determinants of gene expression in diverse cellular and disease states
收藏资源简介:
Code and data for the following publication: Decoding sequence determinants of gene expression in diverse cellular and disease states Avantika Lal*,1, Alexander Karollus*,1,2,3, Laura Gunsalus1, David Garfield4, Surag Nair1, Alex M Tseng1, M Grace Gordon5, John Blischak6, Bryce van de Geijn6, Tushar Bhangale6, Jenna L Collier1, Nathaniel Diamant1, Tommaso Biancalani1, Hector Corrada Bravo1, Gabriele Scalia1, Gokcen Eraslan1 *Equal contributions 1Biology Research | AI Development, gRED Computational Sciences, Genentech, South San Francisco, CA 94080, USA 2School of Computation, Information and Technology, Technical University of Munich, Germany 3Munich Center for Machine Learning 4OMNI Bioinformatics and Department of Regenerative Medicine, Genentech, South San Francisco, CA 94080, USA 5 Department of Cellular and Tissue Genomics, Genentech Research and Early Development, Genentech, South San Francisco, CA 94080, USA 6 Department of Human Genetics, Genentech, South San Francisco, CA 94080, USA Correspondence: Avantika Lal (lal.avantika@gene.com), Gokcen Eraslan (eraslan.gokcen@gene.com)This package contains:1. 4 replicate Decima models (rep0.ckpt, rep1.ckpt, rep2.ckpt, rep3.ckpt)2. Decima software v0.1 (decima-v0.1.tar.gz). The latest version of this software is available at https://github.com/Genentech/decima.3. Decima analysis code (decima-applications.tar.gz)4. Data File 1: An .h5ad file containing Decima’s predictions for all 18,457 genes in the training, validation and test sets, along with metadata for the 18,457 genes and 8,856 pseudobulks in the dataset, and the Pearson correlation between measured and predicted expression values for each pseudobulk and each gene.5. Data File 2: A .h5ad file containing Decima’s predicted effect sizes (log fold change in expression between alternate and reference alleles) for all 573 variants that are identified as high-confidence sc-eQTLs in the fine-mapped OneK1K dataset, as well as negative control variants, in all 8,856 pseudobulks.6. Data File 3: A .h5ad file containing Decima’s predicted effect sizes (log fold change in expression between alternate and reference alleles) for all 837 variants that are identified as high-confidence fine-mapped GWAS causal variants, in all 201 cell types.
本配套代码与数据对应如下发表论文: 解析多样细胞与疾病状态下基因表达的序列决定因素 作者:Avantika Lal*1,Alexander Karollus*1,2,3,Laura Gunsalus1,David Garfield4,Surag Nair1,Alex M Tseng1,M Grace Gordon5,John Blischak6,Bryce van de Geijn6,Tushar Bhangale6,Jenna L Collier1,Nathaniel Diamant1,Tommaso Biancalani1,Hector Corrada Bravo1,Gabriele Scalia1,Gokcen Eraslan1 *共同第一作者 1. 美国加利福尼亚州南旧金山基因泰克(Genentech)gRED计算科学部,生物学研究与人工智能开发 2. 德国慕尼黑工业大学计算、信息与技术学院 3. 慕尼黑机器学习中心 4. 美国加利福尼亚州南旧金山基因泰克OMNI生物信息学与再生医学系 5. 美国加利福尼亚州南旧金山基因泰克研究与早期开发部细胞与组织基因组学系 6. 美国加利福尼亚州南旧金山基因泰克人类遗传学系 通讯作者:Avantika Lal(邮箱:lal.avantika@gene.com)、Gokcen Eraslan(邮箱:eraslan.gokcen@gene.com) 本数据包包含以下内容: 1. 4个重复的Decima模型文件(rep0.ckpt、rep1.ckpt、rep2.ckpt、rep3.ckpt) 2. Decima软件v0.1版本(decima-v0.1.tar.gz),该软件的最新版本可通过https://github.com/Genentech/decima获取 3. Decima分析代码包(decima-applications.tar.gz) 4. 数据文件1:.h5ad格式文件,包含Decima对训练集、验证集与测试集中全部18457个基因的预测结果,同时附带该18457个基因以及数据集中8856个伪bulk样本的元数据,以及每个伪bulk样本与每个基因的实测表达值与预测表达值间的Pearson相关系数 5. 数据文件2:.h5ad格式文件,包含Decima针对精细定位的OneK1K数据集中被鉴定为高可信度单细胞表达数量性状基因座(single-cell expression Quantitative Trait Locus,sc-eQTL)的全部573个变异,以及阴性对照变异,在全部8856个伪bulk样本中的预测效应量(等位基因替代型与参考型之间的表达量对数倍数变化) 6. 数据文件3:.h5ad格式文件,包含Decima针对全部837个被鉴定为高可信度精细定位全基因组关联研究(Genome-Wide Association Study,GWAS)因果变异的位点,在全部201种细胞类型中的预测效应量(等位基因替代型与参考型之间的表达量对数倍数变化)



