VIPER: Training Data and Reproducible Example Dataset for Variant Prioritization
收藏资源简介:
This repository provides data resources to support the reproducibility of the VIPER framework for phenotype-guided variant prioritization. The training data included in this repository are curated from publicly available studies and resources. Detailed descriptions of the original data sources are provided in the manuscript. The data released here are processed versions in JSONL format, including variant annotation, preliminary genotype filtering, and prompt-based feature construction, and are intended for direct input into the model. In addition, a toy example dataset is provided in the same format to demonstrate the expected inputs and enable reviewers and users to run the full pipeline and reproduce representative outputs. Due to ethical and privacy considerations, the real patient-level test datasets used in this study cannot be publicly shared. However, all datasets are available from their original sources or upon appropriate request, subject to relevant approvals. Detailed instructions for data usage and pipeline execution are provided in the repository: https://github.com/VincentLHH/VIPER.



