遇见数据集

Bacterial Training Dataset For Galaxy Training Network Tutorials On Genome Assembly

收藏
Zenodo2020-09-18 更新2026-05-28 收录
数据链接:
官方服务:

资源简介:

This training dataset is from an imaginary <em>Staphylococcus aureus</em> bacterium with a miniature genome. There is a reference genome in various formats as well as some fastq reads of a closely related but also imaginary mutant strain. It is a useful dataset for demonstrating: de novo genome assembly read mapping and variant calling genome annotation The files included are: <strong>wildtype.fna</strong>: the reference genome sequence of the wildtype strain in fasta format (a header line, then the nucleotide sequence of the genome.) <strong>wildtype.gff</strong>: the reference genome sequence of the wildtype strain in general feature format (a list of features - one feature per line, then the nucleotide sequence of the genome.) <strong>wildtype.gbk</strong>: the reference genome sequence in genbank format. <strong>mutant_R1.fastq</strong> and <strong>mutant_R2.fastq</strong>: Fastq sequence reads of a closely related mutant strain. The reads are paired-end. Each read is 150 bases long. The number of bases sequenced is equivalent to 19x the genome sequence of the wildtype strain. (Read coverage 19x - rather low!).

提供机构:
Zenodo
创建时间:
2017-05-23
二维码
社区交流群
二维码
科研交流群
商业服务