Protein coding sequences (CDS) of the genome of a <em>Mus musculus</em>
收藏资源简介:
This dataset contains 22,759 non-redundant protein-coding sequences (CDS) from the Mus musculus C57BL/6 reference genome (GenBank accession GCF_000001635.27_GRCm39). Each CDS is annotated with a gene symbol, GenBank reference ID, and sequence length in base pairs. This curated CDS reference file enables consistent cross-species comparisons at the transcript level and was used to identify orthologous expression patterns and differential regulation in response to viral infection. The values in this dataset are gene-level annotations, GenBank protein references, and sequence sizes. The file is structured as a single spreadsheet, with each row representing a unique gene and its corresponding CDS. The dataset is reusable for any study requiring canonical CDS references from M. musculus C57BL/6, particularly in contexts of comparative transcriptomics. , , # README for CDS_Mus_singles_names.xlsx ## Title: Reference coding sequences from *Mus musculus* (C57BL/6) genome assembly ## Organism: *Mus musculus* (C57BL/6 strain) ## Reference Genome: GenBank accession: GCF_000001635.27_GRCm39 ## File: CDS_Mus_singles_names.xlsx This file contains a curated list of 22,759 non-redundant protein-coding sequences (CDS) from the *Mus musculus* C57BL/6 genome assembly. Each row in the dataset represents a unique gene and includes the gene symbol, corresponding GenBank or UniProt reference, and the coding sequence length in base pairs. ## Contents: CDS_Mus_singles_names.xlsx ## Variables (Columns): 1. gene: Gene symbol (string) 2. reference: GenBank or UniProt protein accession identifier (string) 3. size: Length of the coding sequence in base pairs (integer) ## Units: * Size is measured in base pairs (bp) ## Recommended Citation: If using this dataset, please cite: Genome Reference Consortium Mouse Build 39 (GRCm39), GenBank accession GC..., ,
本数据集包含源自小家鼠(Mus musculus)C57BL/6参考基因组(GenBank登录号GCF_000001635.27_GRCm39)的22759条非冗余蛋白质编码序列(CDS)。每条CDS均注释有基因符号、GenBank参考标识及碱基对长度。该经过人工整理的CDS参考文件可实现转录本水平的跨物种一致性比较,曾被用于鉴定病毒感染响应下的同源表达模式与差异调控机制。本数据集包含基因水平注释、GenBank蛋白质参考信息及序列长度数据,以单一电子表格形式组织,每一行对应一个独特基因及其对应的CDS。该数据集可复用于任何需要小家鼠C57BL/6经典CDS参考序列的研究,尤其适用于比较转录组学相关场景。 # CDS_Mus_singles_names.xlsx 说明文档 ## 标题:小家鼠(C57BL/6品系)基因组组装的经典编码序列参考集 ## 来源生物:小家鼠(Mus musculus,C57BL/6品系) ## 参考基因组:GenBank登录号:GCF_000001635.27_GRCm39 ## 数据文件:CDS_Mus_singles_names.xlsx 本文件包含源自小家鼠C57BL/6基因组组装的22759条非冗余蛋白质编码序列(CDS)整理列表。数据集中每一行对应一个独特基因,包含基因符号、对应的GenBank或UniProt参考标识及编码序列的碱基对长度。 ## 数据集内容:CDS_Mus_singles_names.xlsx ## 变量(列): 1. gene:基因符号(字符串类型) 2. reference:GenBank或UniProt蛋白质登录号标识(字符串类型) 3. size:编码序列碱基对长度(整数类型) ## 单位: * 长度单位为碱基对(bp) ## 推荐引用格式: 若使用本数据集,请引用:基因组参考联盟小鼠组装版本39(GRCm39),GenBank登录号GC...



