遇见数据集

Pangolin precomputed scores

收藏
Zenodo2025-07-16 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains Pangolin precomputed scores for all SNVs in protein-coding genes (hg38 genome version) computed with default parameters: window size 50 nt, scores are masked based on GENCODE splice site annotations (see below for more information, and see the original paper Zeng & Li, 2022). Pangolin is a deep learning model that predicts the effect of a variant on the splice site usage. It computes a gain and a loss score for every position within a user-defined window around the variant that represents the increase and decrease in the usage of a potential splice site at the respective positions. Pangolin outputs the maximum gain and the maximum loss scores within the window together with the corresponding positions. It also provides an option to mask scores when a genome annotation is provided to the model, which sets those scores to zero if Pangolin predicts activation for annotated splice sites and deactivation for unannotated splice sites. The directory contains per-gene TSV files with the following columns: chrom: Chromosome pos: Genomic position ref: Reference allele alt: Alternative allele gain_score: Pangolin gain score gain_pos: relative position of the gain score loss_score: Pangolin loss score loss_pos: relative position of the loss score

本数据集包含hg38基因组版本下蛋白质编码基因中所有单核苷酸变异(SNVs)的Pangolin预计算分值,计算时采用默认参数:窗口大小为50 nt,分值基于GENCODE剪接位点注释进行掩码(详细说明见下文,原始论文参见Zeng与Li, 2022)。 Pangolin是一款用于预测变异对剪接位点使用影响的深度学习模型。它会在变异侧翼的用户自定义窗口内,为每个位置计算增益分值与损失分值,分别代表对应位置潜在剪接位点的使用量上升与下降幅度。Pangolin会输出该窗口内的最大增益分值与最大损失分值,以及其对应的相对位置。此外,当向模型传入基因组注释信息时,模型支持对分值进行掩码:若Pangolin预测注释剪接位点被激活、未注释剪接位点被失活,则将对应分值置零。 该数据集目录包含按基因拆分的TSV文件,各文件包含以下列: chrom: 染色体(Chromosome) pos: 基因组位置(Genomic position) ref: 参考等位基因(Reference allele) alt: 替代等位基因(Alternative allele) gain_score: Pangolin增益分值 gain_pos: 增益分值的相对位置 loss_score: Pangolin损失分值 loss_pos: 损失分值的相对位置

提供机构:
Zenodo
创建时间:
2025-07-16
二维码
社区交流群
二维码
科研交流群
商业服务