A New Comprehensive Annotation of Leucine-Rich Repeat-Containing Receptors in Rice
收藏资源简介:
Datasets for preprint (https://doi.org/10.1101/2021.01.29.428842) entitled "<strong>A new comprehensive Annotation of Leucine-Rich Repeat-Containing Receptors in Rice</strong>". This paper describes an in-depth manual curation of LRR-CR annotations including genes containing nonsense mutations, tagged as 'non-canonical', by opposition to 'canonical', that have expected gene models. Contains 7 files for each rice cultivar: - domain annotation for each LRR-CR protein (xxx_LRR_domains_filtered.csv) - gff file containing LRR-CR gene model annotations - fasta file (1): the complete genes 'gene' (nucleotide sequence, exons and introns) - fasta file (2): the coding sequences 'CDS' (nucleotide sequence, exons, translatable). Note that for genes experiencing frameshift, the one or two bases that cause the frameshift are avoided in order to have nucleotide sequence that can be translated. Terminal and in frame stop codons are encoded by a '*'. - fasta file (3): protein sequences 'PEP' (amino acid sequence, exons, corresponding to translation of CDS (2) ) - fasta file (4): 'cDNA' (nucleotide sequence, exons). Note that for genes experiencing frameshift, the translated protein sequence will not correspond to the one present in the file (3). For all other genes, the files "CDS" and "cDNA" retrieve the same nucleotide sequences. - fasta file (5): 'cDNA_wFrameshit' (nucleotide sequence, exons) : same sequences than in (4) except for genes experiencing frameshift whose sequences are completed with one or two "!" characters at the position of the frameshift in order to conserve the right reading frame (also used by V. Ranwez et al. for the MACSE programs https://dx.doi.org/10.1093/molbev/msy159 ).
本数据集对应预印本论文(https://doi.org/10.1101/2021.01.29.428842),题为**《水稻中富含亮氨酸重复序列受体的全新综合注释》**。该论文针对富含亮氨酸重复序列受体(Leucine-Rich Repeat-Containing Receptors, LRR-CR)的注释开展了深度人工手动校验,涵盖携带无义突变、被标记为"非经典(non-canonical)"的基因,与具备标准基因模型的"经典(canonical)"基因形成对照。每个水稻品种对应7个数据文件: - 各LRR-CR蛋白的结构域注释文件(xxx_LRR_domains_filtered.csv) - 包含LRR-CR基因模型注释的通用特征格式(General Feature Format, GFF)文件 - 快速序列比对格式(FASTA)文件1:完整基因序列(gene),即包含外显子与内含子的核苷酸序列 - 快速序列比对格式文件2:编码序列(Coding Sequence, CDS),即可翻译的外显子核苷酸序列。针对存在移码突变的基因,会规避1或2个导致移码的碱基,以确保核苷酸序列可正常翻译;终端且处于读码框内的终止密码子以字符`*`表示 - 快速序列比对格式文件3:蛋白质序列(PEP),即由上述CDS(文件2)翻译得到的、对应外显子区域的氨基酸序列 - 快速序列比对格式文件4:互补DNA序列(Complementary DNA, cDNA),即外显子区域的核苷酸序列。注意:对于存在移码突变的基因,该文件翻译得到的蛋白质序列与文件3的蛋白质序列并不一致;对于其余基因,CDS文件与cDNA文件的核苷酸序列完全相同 - 快速序列比对格式文件5:带移码标记的互补DNA序列(cDNA_wFrameshift):与文件4的序列基本一致,仅针对存在移码突变的基因,会在移码位置插入1或2个`!`字符以维持正确的读码框(该标记方法被V. Ranwez等人应用于MACSE程序,详见https://dx.doi.org/10.1093/molbev/msy159)



