HAVEN Datasets
收藏官方服务:
资源简介:
This folder contains the datasets used for pre-training, fine-tuning, and evaluating HAVEN: Hierarchical Attention for Viral protEin-based host iNference. HAVEN is a viral protein language model pre-trained on 1.2 million protein sequences belonging to Viridae (viruses). It is fine-tuned to predict the hosts from which the viral protein sequence was sampled ("virus-host"). The viral protein sequences are downloaded from UniRef90. The host for the viral sequences are from the European Nucleotide Archive (ENA) maintained by European Molecular Biology Laboratory-European Bioinformatics Institute (EMBL-EBI).
提供机构:
Zenodo创建时间:
2025-05-28



