遇见数据集

Pangolin precomputed scores

收藏
Zenodo2025-07-16 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains Pangolin precomputed scores for all SNVs in protein-coding genes (hg38 genome version) computed with default parameters: window size 50 nt, scores are masked based on GENCODE splice site annotations (see below for more information, and see the original paper Zeng & Li, 2022). Pangolin is a deep learning model that predicts the effect of a variant on the splice site usage. It computes a gain and a loss score for every position within a user-defined window around the variant that represents the increase and decrease in the usage of a potential splice site at the respective positions. Pangolin outputs the maximum gain and the maximum loss scores within the window together with the corresponding positions. It also provides an option to mask scores when a genome annotation is provided to the model, which sets those scores to zero if Pangolin predicts activation for annotated splice sites and deactivation for unannotated splice sites. The directory contains per-gene TSV files with the following columns: chrom: Chromosome pos: Genomic position ref: Reference allele alt: Alternative allele gain_score: Pangolin gain score gain_pos: relative position of the gain score loss_score: Pangolin loss score loss_pos: relative position of the loss score

提供机构:
Zenodo
创建时间:
2025-07-16
二维码
社区交流群
二维码
科研交流群
商业服务