遇见数据集

An Arabidopsis gene expression matrix derived from large-scaled RNA-Seq datasets

收藏
Figshare2021-05-06 更新2026-04-28 收录
官方服务:

资源简介:

The dataset contains a gene expression matrix to be used with the EXPLICIT package to construct an gene expression predictor. It has 24545 RNA-Seq samples with 38194 genes in total. Two versions of the matrix are provided. It is recommended to use the smaller version first (At.matrix.demo.h5) to get familiar with the EXPLICIT package, and then change to the full version (At.matrix.full.h5). The smaller matrix contains 5000 samples randomly selected from all 24545 samples.At.matrix.demo.h5|--expression_log2cpm -------- (5000 samples [row] X 38194 genes [col])|--gene_name ------------------ ( for the 38194 genes)|--rnaseq_id ------------------- (for the 5000 samples)|--idx_tf_gene ------- (specifying TF genes used for model construction)|--idx_target_gene --- (specifying target genes used for model construction)|--independent_samples_for_validation >>>|--expression_log2cpm ------ (2 samples [row] X 38194 genes [col]) >>>|--gene_name ----------------- (for the 38194 genes) >>>|--sample_id ---------------- (for the 2 RNA-Seq samples)At.matrix.full.h5|--expression_log2cpm -------- (24545 samples [row] X 38194 genes [col])|--gene_name ------------------ (for the 38194 genes)|--rnaseq_id ------------------- (for the 24545 samples)|--idx_tf_gene ------- (specifying TF genes used for model construction)|--idx_target_gene --- (specifying target genes used for model construction)|--independent_samples_for_validation >>>|--expression_log2cpm ------ (2 samples [row] X 38194 genes [col]) >>>|--gene_name ----------------- (for the 38194 genes) >>>|--sample_id ---------------- (for the 2 RNA-Seq samples)

本数据集包含基因表达矩阵,可配合EXPLICIT软件包构建基因表达预测器。该数据集共计24545个RNA测序(RNA-Seq)样本,涵盖38194个基因。本次发布共提供该矩阵的两个版本,建议优先使用较小版本(At.matrix.demo.h5)以熟悉EXPLICIT软件包,后续再切换至完整版本(At.matrix.full.h5)。该较小矩阵是从全部24545个样本中随机抽取的5000个样本构建而成。 At.matrix.demo.h5|--expression_log2cpm -------- (5000个样本[行] × 38194个基因[列])|--gene_name ------------------ (对应38194个基因的名称)|--rnaseq_id ------------------- (对应5000个样本的RNA测序编号)|--idx_tf_gene ------- (用于指定模型构建所用的转录因子(Transcription Factor, TF)基因)|--idx_target_gene --- (用于指定模型构建所用的靶基因)|--independent_samples_for_validation >>>|--expression_log2cpm ------ (2个样本[行] × 38194个基因[列]) >>>|--gene_name ----------------- (对应38194个基因的名称) >>>|--sample_id ---------------- (对应2个RNA测序样本的编号) At.matrix.full.h5|--expression_log2cpm -------- (24545个样本[行] × 38194个基因[列])|--gene_name ------------------ (对应38194个基因的名称)|--rnaseq_id ------------------- (对应24545个样本的RNA测序编号)|--idx_tf_gene ------- (用于指定模型构建所用的转录因子(Transcription Factor, TF)基因)|--idx_target_gene --- (用于指定模型构建所用的靶基因)|--independent_samples_for_validation >>>|--expression_log2cpm ------ (2个样本[行] × 38194个基因[列]) >>>|--gene_name ----------------- (对应38194个基因的名称) >>>|--sample_id ---------------- (对应2个RNA测序样本的编号)

创建时间:
2021-05-06
二维码
社区交流群
二维码
科研交流群
商业服务