Surface Immune Signalling Unlocks NLR Activation Through mRNA Alternative Splicing
收藏资源简介:
This dataset contains the sequence resources, motif-search outputs, annotation tables, and phylogenetic results used in the Pi-Alternative-Splicing-Project for plant NLR analysis, with particular emphasis on N-terminal MADA motif characterization and downstream comparative analysis. The archive includes four main components: 1. HMM-based MADA motif search resourcesThis part contains the MADA profile HMM (MADA.hmm), the reference NLR protein sequence set used as the search database (NLR_ref.fasta; 74,158 sequences), and the corresponding HMMER output files, including full output (MADA.out), per-sequence table output (MADA_tbl.out), and per-domain table output (MADA_domtbl.out). These files document the identification of candidate MADA-containing NLR proteins. 2. MADA extension and metadata tablesThis component contains curated sequence sets and metadata tables for MADA-positive candidates and related genomic extractions. Included files comprise:- amino acid sequences of MADA-containing NLRs (SolDB_MADA_ext.fasta; 619 records),- NB-ARC domain sequences of the same candidates (SolDB_MADA_ext_NBARC.fasta; 619 records),- candidate ID lists,- annotation and summary tables,- metadata-integrated tables,- genomic coordinate tables, and- strand-aware genomic sequence extractions with flanking regions (SolDB_MADA_ext_redundant_genomics.fasta; 1,642 records).These files support candidate tracking, motif annotation, coordinate matching, and downstream comparative analyses. 3. Sequence collections for NLR filtering and reference buildingThis component includes large-scale filtered NLR protein datasets and reference sequence resources:- NLR_filtered_seq.fasta (169,434 sequences),- clustered non-redundant representative sequences NLR_filtered_seq_clu.fasta (87,460 sequences),- corresponding CD-HIT clustering assignments (.clstr),- RefPlantNLR.fasta (481 curated reference plant NLR proteins),- RefPlantNLR_NBARC.fasta (406 curated reference NB-ARC sequences), and- NBARC_ref.fasta (74,083 extracted NB-ARC domain sequences).These files were used for sequence filtering, reference comparison, redundancy reduction, and domain-level downstream analysis. 4. Phylogenetic resultThe archive also includes a rooted phylogenetic tree based on NB-ARC reference/domain sequences (NBARC_ref_famsa_rooted.tree), used for comparative evolutionary analysis of candidate NLRs. Together, these files provide the supporting data resources for motif detection, NLR curation, sequence redundancy filtering, genomic coordinate integration, and phylogenetic interpretation in this project. File formats include FASTA, CSV, TXT, HMMER output tables, CD-HIT cluster files, and Newick tree files.
本数据集涵盖了Pi可变剪接项目中用于植物NLR(NLR)分析的序列资源、基序(motif)搜索结果、注释表格以及系统发育分析结果,重点聚焦于N端MADA基序(motif)的特征解析与后续比较分析。 该归档文件包含四大核心组成部分: 1. 基于隐马尔可夫模型(Hidden Markov Model)的MADA基序(motif)搜索资源 本部分包含MADA隐马尔可夫模型配置文件(MADA.hmm)、用作搜索数据库的参考NLR蛋白序列集(NLR_ref.fasta,共74158条序列),以及对应的HMMER(HMMER)输出文件,包括全量输出文件(MADA.out)、单序列表格输出文件(MADA_tbl.out)与单结构域表格输出文件(MADA_domtbl.out)。上述文件用于记录含MADA基序(motif)的候选NLR蛋白的鉴定结果。 2. MADA延伸序列与元数据表 本组件包含经人工整理的MADA阳性候选蛋白序列集与元数据表,以及相关基因组提取数据。包含的文件如下: - 含MADA基序(motif)的NLR蛋白氨基酸序列文件(SolDB_MADA_ext.fasta,共619条记录) - 上述候选蛋白的NB-ARC结构域(NB-ARC)序列文件(SolDB_MADA_ext_NBARC.fasta,共619条记录) - 候选蛋白ID列表 - 注释与汇总表格 - 整合元数据表 - 基因组坐标表格 - 带链方向信息的侧翼区基因组序列提取文件(SolDB_MADA_ext_redundant_genomics.fasta,共1642条记录) 上述文件可用于候选蛋白追踪、基序(motif)注释、坐标匹配及后续比较分析。 3. 用于NLR筛选与参考集构建的序列集合 本组件包含大规模经筛选的NLR蛋白数据集与参考序列资源: - NLR_filtered_seq.fasta(共169434条序列) - 聚类去重代表序列文件NLR_filtered_seq_clu.fasta(共87460条序列) - 对应的CD-HIT(CD-HIT)聚类分配文件(.clstr) - RefPlantNLR.fasta(共481条经人工整理的植物NLR参考蛋白序列) - RefPlantNLR_NBARC.fasta(共406条经人工整理的NB-ARC结构域参考序列) - NBARC_ref.fasta(共74083条提取得到的NB-ARC结构域序列) 上述文件用于序列筛选、参考序列比对、冗余去除及结构域层面的后续分析。 4. 系统发育分析结果 本归档文件还包含基于NB-ARC结构域参考序列构建的有根系统发育树文件(NBARC_ref_famsa_rooted.tree),用于候选NLR蛋白的比较进化分析。 综上,本数据集的所有文件可为该项目中的基序(motif)检测、NLR蛋白人工整理、序列冗余过滤、基因组坐标整合及系统发育结果解读提供支撑数据资源。 本数据集包含的文件格式包括FASTA格式(FASTA)、CSV格式(CSV)、TXT格式(TXT)、HMMER(HMMER)输出表格、CD-HIT(CD-HIT)聚类文件以及Newick树文件(Newick)。



