Human open reading frame data
收藏资源简介:
Data used in the human de novo ORF paper (Dowling et al. 2020 Genome Biology and Evolution evaa194). The AA and DNA sequences were used to predict sequence properties such as aggregation propensity and intrinsic structural disorder.<br> In the publication figures 1, 3, 4, 5, and 7 are based on this data.<br> Figure 2 uses different ORFs and figure 6 uses these ORFs and their predicted chimpanzee homologs<br> Last updated: 24/09/2020 Contains:<br> 1. hsapiens.orfs.aa.fa <br> - all human ORFs as amino-acid sequences<br> - contains 36524 sequences<br> Note: this has not been filtered for minimum expression of 0.5 TPM<br> <br> 2. hsapiens.orfs.dna.fa <br> - all human ORFs as DNA seqeunces<br> - contains 36524 sequences<br> Note: this has not been filtered for minimum expression of 0.5 TPM<br> <br> 3. orf_age_annotation.csv<br> - The 29751 ORFs used for the plots in the human de novo publication.<br> - Contains ORF_id, minimum age (million years) and annotation status.<br> - ORFs here where filtered for a minimum expression value of 0.5 TPM<br> Note: In the publication figures were made using ORFs belong to annotation classes 0 (intergenic), 3 (intron), and 5,6, and 7 (CDS/exon).<br> 0 = intergenic<br> 1 = close to gene (same strand)<br> 2 = close to gene (opposite strand)<br> 3 = intron (same strand)<br> 4 = intron (opposite strand)<br> 5 = CDS (out of frame, same strand)<br> 6 = CDS (out of frame, opposite strand)<br> 7 = CDS (in frame)<br> - information on minimum age given in million years (column 'age') <br> 0 = human-restricted<br> 6.65 = homolog in Pan<br> 9.06 = homolog in Gorilla<br> 15.76 = homolog in Pongo<br> 28.44 = homolog in Macaca<br> 90 = homolog in Mus<br> Note: in the publication these correspond to conservation classes 0-5 respectively.<br>



