遇见数据集

Fastq files

收藏
Mendeley Data2024-01-31 更新2024-06-28 收录
官方服务:

资源简介:

The project contains the fastq files (in a gzipped format) from simulated RNA-seq data for 12 paired-ended samples (i.e., biological replicates); all reads are 101 base pairs long.The first subscript denotes the sample id (1 to 12), while the second subscript indicates the two strands of each sample (1 or 2).Samples belong to two groups: samples 1 to 6 constitute the first group, while samples 7 to 12 represent the second group.In each group 1,000 genes exhibit differential transcript usage (DTU) and further 1,000 genes (partially overlapping with the previous set) shows differential gene expression (DGE) between groups. Genes showing DTU were simulated by randomly permuting the relative abundance of the four most expressed transcripts; for gene with two or three transcripts only, all transcripts relative abundances were permuted. The dataset is used to benchmark DTU methods in the manuscript entitled "BANDITS: Bayesian differential splicing accounting for sample-to-sample variability and mapping uncertainty". The code for simulating the data is available on GitHub at https://github.com/SimoneTiberi/BANDITS_manuscript .

本数据集包含12个双端(paired-ended)测序的生物学重复样本的模拟RNA测序(RNA-seq)数据的FASTQ文件,所有文件均采用gzip压缩格式;所有测序读段长度均为101个碱基对。第一个下标代表样本ID(1至12),第二个下标则表示每个样本的两条测序读段(1或2)。所有样本分为两组:样本1至6为第一组,样本7至12为第二组。两组间存在1000个基因表现出转录本使用差异(differential transcript usage, DTU),另有1000个基因(与前述基因集部分重叠)表现出基因表达差异(differential gene expression, DGE)。表现出DTU的基因通过随机置换表达量最高的四种转录本的相对丰度进行模拟;对于仅含有2或3种转录本的基因,则对其所有转录本的相对丰度进行置换。本数据集用于在题为《BANDITS:考虑样本间变异与比对不确定性的贝叶斯差异剪接分析》的手稿中对DTU分析方法进行基准测试。该模拟数据的生成代码可在GitHub平台获取,链接为https://github.com/SimoneTiberi/BANDITS_manuscript。

创建时间:
2024-01-31
二维码
社区交流群
二维码
科研交流群
商业服务