PacBio Iso-Seq, transcript models
收藏资源简介:
Structural annotation/transcript models of RNA (cDNA) sequenced from six different tissues (brain, ileum, lung, ovary, spleen, testis) from the tufted duck on PacBio platform.Full-length non-chimeric (FLNC) reads were built according to PacBio's IsoSeq3 pipeline (ccs, lima, refine), fasta sequences extracted with Bamtools and poly-A tails trimmed with the TAMA tool tama_flnc_polya_cleanup.py. The FLNC reads were mapped to the tufted duck reference genome with Minimap2 and redundant transcript models collapsed with the TAMA tool tama_collapse.py.ccs options: ${--noPolish --minLength=300 --minPasses=1 --minZScore=-999 --maxDropFraction=0.8 --minPredictedAccuracy=0.8 --minSnr=4}lima options: ${--isoseq --dump-clips --no-pbi -j 4 Teloprime.fasta}refine options: ${Teloprime.fasta}Minimap2 options: ${--secondary=no -a -x splice -u f -C 5}tama_collapse.py options: ${-c 95 -x capped -a 100 -z 100}
本数据集包含从凤头鸭(tufted duck)的6种不同组织(脑、回肠、肺、卵巢、脾脏、睾丸)中提取的RNA(cDNA,互补DNA)的测序数据对应的结构注释与转录本模型,所有测序工作均基于PacBio测序平台(Pacific Biosciences)完成。 依据PacBio官方的IsoSeq3分析流程(包含ccs、lima、refine三个子步骤)构建全长非嵌合读段(Full-length non-chimeric,FLNC),使用Bamtools工具提取FASTA序列,并通过TAMA工具的tama_flnc_polya_cleanup.py脚本切除poly-A尾。 使用Minimap2工具将FLNC读段比对至凤头鸭参考基因组,并通过TAMA工具的tama_collapse.py脚本对冗余转录本模型进行合并去重。 ccs参数设置:${--noPolish --minLength=300 --minPasses=1 --minZScore=-999 --maxDropFraction=0.8 --minPredictedAccuracy=0.8 --minSnr=4} lima参数设置:${--isoseq --dump-clips --no-pbi -j 4 Teloprime.fasta} refine参数设置:${Teloprime.fasta} Minimap2参数设置:${--secondary=no -a -x splice -u f -C 5} tama_collapse.py参数设置:${-c 95 -x capped -a 100 -z 100}



