RefSeq Data & Scripts for ORF dominance
收藏资源简介:
This is the fileset for the article "Protein-coding potential of RNAs measured by open reading frame dominance" by Y.Suenaga, et al. This fileset consists of the datasets of human (Feb 2015 and April 2018) and 8 spicies given in the article. [Dataset] - RefSeq gzipped fasta files (data/RefSeq) - Scripts to generate open reading frame (ORF) dominance score and other information from RefSeq data. (script) [How to Run Scripts] - 1. Unpack a tar.xz file. - 2. Run script/01_MergeFa.sh. - 3. Run script/02_ORFdominance.sh. - 4. Run script/03_Format.sh Under data/03_Format, you can get NM.txt and NR.txt as the result. [Note] - Scripts require Linux, bash, and perl. - Scripts use randomized data. Consequently, results differ slightly for the same input data. - Scripts are the same files in "Scripts for ORF dominance" (DOI: 10.6084/m9.figshare.7269518).



