Surface Immune Signalling Unlocks NLR Activation Through mRNA Alternative Splicing
收藏资源简介:
This dataset contains the sequence resources, motif-search outputs, annotation tables, and phylogenetic results used in the Pi-Alternative-Splicing-Project for plant NLR analysis, with particular emphasis on N-terminal MADA motif characterization and downstream comparative analysis. The archive includes four main components: 1. HMM-based MADA motif search resourcesThis part contains the MADA profile HMM (MADA.hmm), the reference NLR protein sequence set used as the search database (NLR_ref.fasta; 74,158 sequences), and the corresponding HMMER output files, including full output (MADA.out), per-sequence table output (MADA_tbl.out), and per-domain table output (MADA_domtbl.out). These files document the identification of candidate MADA-containing NLR proteins. 2. MADA extension and metadata tablesThis component contains curated sequence sets and metadata tables for MADA-positive candidates and related genomic extractions. Included files comprise:- amino acid sequences of MADA-containing NLRs (SolDB_MADA_ext.fasta; 619 records),- NB-ARC domain sequences of the same candidates (SolDB_MADA_ext_NBARC.fasta; 619 records),- candidate ID lists,- annotation and summary tables,- metadata-integrated tables,- genomic coordinate tables, and- strand-aware genomic sequence extractions with flanking regions (SolDB_MADA_ext_redundant_genomics.fasta; 1,642 records).These files support candidate tracking, motif annotation, coordinate matching, and downstream comparative analyses. 3. Sequence collections for NLR filtering and reference buildingThis component includes large-scale filtered NLR protein datasets and reference sequence resources:- NLR_filtered_seq.fasta (169,434 sequences),- clustered non-redundant representative sequences NLR_filtered_seq_clu.fasta (87,460 sequences),- corresponding CD-HIT clustering assignments (.clstr),- RefPlantNLR.fasta (481 curated reference plant NLR proteins),- RefPlantNLR_NBARC.fasta (406 curated reference NB-ARC sequences), and- NBARC_ref.fasta (74,083 extracted NB-ARC domain sequences).These files were used for sequence filtering, reference comparison, redundancy reduction, and domain-level downstream analysis. 4. Phylogenetic resultThe archive also includes a rooted phylogenetic tree based on NB-ARC reference/domain sequences (NBARC_ref_famsa_rooted.tree), used for comparative evolutionary analysis of candidate NLRs. Together, these files provide the supporting data resources for motif detection, NLR curation, sequence redundancy filtering, genomic coordinate integration, and phylogenetic interpretation in this project. File formats include FASTA, CSV, TXT, HMMER output tables, CD-HIT cluster files, and Newick tree files.



