UltraSynBERT_All_Archives
收藏资源简介:
Decoding the RNA interactome by UltraSynBERT The official repository for the article "Decoding the RNA interactome by UltraSynBERT". Overview RNA encodes complex molecular behaviors that underlie binding, regulation, and cellular fate. Current RNA language models have been shaped primarily by endogenous transcripts, leaving the potential of exogenous RNAs to extend model boundaries unexplored. We present UltraSynBERT, an RNA language model pretrained on synthetic RNA binding landscapes to capture RNA binding properties across in vitro and in vivo contexts. Fine-tuned on in vitro SELEX datasets, the model identifies RNA aptamers for diverse targets, including small molecules, proteins, cells, and tissues. The model also predicts tissue specificity for millions of RNA species across 22 human organs based on their 3’-UTR sequences, captures humanpathogenic viral RNA tropism, recognizes RNA base modifications, and characterizes SARS-CoV-2 replicase binding with single-base resolution features. UltraSynBERT Archives 1. Pre-training datasets and models In this repository, you will find the following pre-training datasets and pre-trained model checkpoints: Model Name Parameters Pre-training Data RNA Species Number Description Checkpoint UltraSynBERT 33.5M UltraSelex SiR-B 10 million pre-training from scratch UltraSynBERT.pkl UltraSynBERTsource_SiR 33.5M SELEX SiR 10 million pre-training from scratch UltraSynBERT_source_SiR.pkl UltraSynBERTsource-nsp12 33.5M SELEX nsp12 10 million pre-training from scratch UltraSynBERT_source_nsp12.pkl UltraSynBERTmolecules 33.5M 12 SELEX targets’ training sets 3.24 million continued pre-training the basic UltraSynBERT model UltraSynBERT_molecule.pkl UltraSynBERT3UTR 33.5M The preliminary 3’-UTR full training datasets 1.69 million continued pre-training the basic UltraSynBERT model UltraSynBERT_3UTR.pkl UltraSynBERTplus 33.5M 12 SELEX targets’ training sets and the preliminary 3’-UTR full training datasets 4.93 million continued pre-training the basic UltraSynBERT model UltraSynBERT_plus.pkl UltraSynBERTRIP 33.5M Human RIP-Seq in ENCODE 6.88 million continued pre-training the basic UltraSynBERT model UltraSynBERT_RIP.pkl NOTE: The pretraining data is located in the Dataset Availability\1. Dataset for pre-training folder. The pretrained models can be found in the Model Availability\1. Pre-training models and Model Availability\2. Continued pre-training models folders. 2. Downstream task datasets and fine-tuned models Additionally, this official repository also provides all datasets for downstream tasks and supervised fine-tuned models based on the base UltraSynBERT for downstream applications, including: RNA aptamer identification models targeting small molecules, proteins, cells, and tissues (12 specialized datasets), and tissue-specificity prediction for human mRNA 3'-UTR regions. Task Type Data Type Dataset Number Dataset SiR-binding prediction In vitro 3 (1) UltraSelex-SiR-B.csv (2) SELEX-SiR-whole.csv (3) SELEX-SiR-exclusive.csv In vitro RNA-target binding prediction In vitro 12 (1) SELEX-small-molecule-DAse.csv; (2) SELEX-small-molecule-BC.csv: (3) SELEX-small-molecule-PR.csv (4) SELEX-small-molecule-MI.csv: (5) SELEX-protein-TARDBP.csv: (6) SELEX-protein-RT.csv (7) SELEX-protein-RBM24.csv; (8) SELEX-protein-S15.csv; (9) SELEX-(multi)cellular-ISLETS.csv (10) SELEX-(multi)cellular-MDSC.csv; (11) SELEX-(multi)cellular-CHO-K1.csv;(12) SELEX-(multi)cellular-TNBC.csv In vivo RNA-protein interactions prediction (human source) In vivo 11 (1) 16_ICLIP_hnRNPC_Hela_iCLIP_all_clusters_sequences.csv; (2) 17_ICLIP_HNRNPC_hg19_sequences.csv;(3) 18_ICLIP_hnRNPL_Hela_group_3975_all-hnRNPL-Hela-hg19_sum_G_hg19--ensembl59_from_2337-2339-741_bedGraph-cDNA-hits-in-genome_sequences.csv (4) 19_ICLIP_hnRNPL_U266_group_3986_all-hnRNPL-U266-hg19_sum_G_hg19--ensembl59_from_2485_bedGraph-cDNA-hits-in-genome_sequences.csv; (5) 20_ICLIP_hnRNPlike_U266_group_4000_all-hnRNPLlike-U266-hg19_sum_G_hg19--ensembl59_from_2342-2486_bedGraph-cDNA-hits-in-genome_sequences.csv; (6) 22_ICLIP_NSUN2_293_group_4007_all-NSUN2-293-hg19_sum_G_hg19--ensembl59_from_3137-3202_bedGraph-cDNA-hits-in-genome_sequences.csv (7) 27_ICLIP_TDP43_hg19_sequences.csv; (8) 28_ICLIP_TIA1_hg19_sequences.csv; (9) 29_ICLIP_TIAL1_hg19_sequences.csv (10) 30_ICLIP_U2AF65_Hela_iCLIP_ctrl_all_clusters_sequences.csv; (11) 31_ICLIP_U2AF65_Hela_iCLIP_ctrl+kd_all_clusters_sequences.csv In vivo RNA-protein interactions prediction (mouse source) In vivo 11 (1) CLIP-mouse-EZH2-sequences.csv; (2) CLIP-mouse-FUS-sequences.csv; (3) CLIP-mouse-HNRNPR-sequences.csv (4) CLIP-mouse-LIN28A-sequences.csv; (5) CLIP-mouse-RBFOX2-sequences.csv; (6) CLIP-mouse-RBM10-sequences.csv (7) CLIP-mouse-SRSF2-sequences.csv; (8) CLIP-mouse-SRSF3-sequences.csv; (9) CLIP-mouse-TARDBP-sequences.csv (10) CLIP-mouse-U2AF2-sequences.csv; (11) CLIP-mouse-YTHDC2-sequences.csv modification In vivo 13 m6A: (1) m6A-A549.csv; (2) m6A-CD8T.csv; (3) m6A-ESC.csv (4) m6A-HCT116.csv; (5) m6A-HEK293.csv; (6) m6A-HEK293T.csv (7) m6A-Hela.csv; (8) m6A-HepG2.csv; (9) m6A-MOLM13.csv m1A: m1A_dataset_with_split.csv m5C: m5c_dataset_with_split.csv m6Am: m6Am_dataset_with_split.csv pseudouridine: pseudouridine_dataset_with_split.csv 3'UTR tissue-specific recognition In vivo 1 human_22_tissue_three_terminal_UTR.csv Zero-shot RNA mutation effect prediction targeting SARs-CoV-2 replicase nsp12 In vitro The wild-type sequence and seven mutated sequences RNA-mutation.fasta RNA post-transcriptional regulation–related tasks in vitro/In vivo 4 (1) miRNA _interactions.csv; (2) mRNA_subcellular_localization.csv; (3) Polyadenylation_signal.csv; (4) RNA_stability.csv human-pathogenic RNA viruses In vivo 1 virus.csv NOTE: The datasets for downstream tasks are in the Dataset Availability\2. Dataset for fine-tuning folder. The fine-tuned models for downstream tasks are available in the Model Availability\3. Fine-tuned models folder.



