SupplementaryDataD1 for: Overcoming the widespread flaws in the annotation of vertebrate selenoprotein genes in public databases
收藏资源简介:
Supplementary Data D1. Selenoprotein predictions by Selenoprofiles. Here's the content of the included README file for a description of the files included. --- ## Directory Overview ### NCBI.selenoprotein_predictions.zip Contains selenoprotein prediction data sourced from the NCBI database. Inside this file you will find a directory for each species with the following format: {species_name}.{accession_id} Inside each directory you will find: - `{species_name}.{accession_id}.sec.011225.gtf` & `{species_name}.{accession_id}.CDS.011225.gtf` GTF file describing genomic intervals of predicted selenoproteins. The relevant information by Selenoprofiles Orthology and Lineage, as well as the misannotation status, is also available in each row. **sec gtf file** contains only selenocysteine prediction residues from each selenoprotein gene **CDS gtf file** contains coding sequence entries for all Selenoprofiles predictions. All those that are not selenoprotein predictions (i.e. Selenoprofiles label is not "selenocysteine") are marked as such. - `{species_name}.{accession_id}.AGGREGATE.011225.tsv` & `{species_name}.{accession_id}.COMPLETE.011225.tsv` TSV file describing the (mis)annotation status of predicted selenoprotein genes. Selenoprofiles Orthology and Lineage information is also included. **COMPLETE tsv file** contains (mis)annotation labels at the transcript level. When multiple annotated transcripts overlap a Selenoprofiles selenoprotein prediction, one label is provided for each. **AGGREGATE tsv file** contains a single (mis)annotation label per selenoprotein prediction (gene level annotation). - `output folder` Folder containing selenoprofiles predictions in two different text formats: **p2g** and **ali** (Selenoprofiles formats, see below and online documentation). - `results.sqlite` SQlite database where selenoprofiles results are saved (Selenoprofiles format, see online documentation). --- ### Ensembl.selenoprotein_predictions.zip Contains selenoprotein prediction data sourced from the Ensembl database. Inside this file you will find a directory for each species with the following format: {species_name} Inside each directory you will find: - `{species_name}.sec.011225.gtf` & `{species_name}.CDS.011225.gtf` GTF file describing genomic intervals of predicted selenoproteins. The relevant information by Selenoprofiles Orthology and Lineage, as well as the misannotation status, is also available in each row. **sec gtf file** contains only selenocysteine prediction residues from each selenoprotein gene **CDS gtf file** contains coding sequence entries for all Selenoprofiles predictions. All those that are not selenoprotein predictions (i.e. Selenoprofiles label is not "selenocysteine") are marked as such. - `{species_name}.AGGREGATE.011225.tsv` & `{species_name}.COMPLETE.011225.tsv` TSV file describing the (mis)annotation status of predicted selenoprotein genes. Selenoprofiles Orthology and Lineage information is also included. **COMPLETE tsv file** contains (mis)annotation labels at the transcript level. When multiple annotated transcripts overlap a Selenoprofiles selenoprotein prediction, one label is provided for each. **AGGREGATE tsv file** contains a single (mis)annotation label per selenoprotein prediction (gene level annotation). - `output folder` Folder containing selenoprofiles predictions in two different text formats: **p2g** and **ali** (Selenoprofiles formats, see below and online documentation). - `results.sqlite` SQlite database where selenoprofiles results are saved (Selenoprofiles format, see online documentation). Besides the species listed in the paper, we also provide files for these non-vertebrate species: *Drosophila melanogaster*, *Ciona intestinalis*, *Ciona Savignyi*, *Caenorhabditis elegans* and *Saccharomyces cerevisiae* --- ## File Descriptions ### FASTA Files (`*.ali`) These files contain the selenoprotein family profile sequences, aligned with the Selenoprofiles predicted selenoprotein sequences.**Header format:** (for Selenoprofiles predictions, not profile sequences) Identifier (first word of each title): profile_name.index.label - `profile_name`: identifies the source profile for the prediction - `index`: an arbitrary numeric ID - `label`: identifies the class of predicted gene (e.g. selenocysteine, cysteine) Rest of fasta title: - `chromosome`: the contig or chromosome identifier of the prediction- `strand`: the strand (+ or -) of the prediction- `positions`: string of coding exons positions (e.g. 61511-61713,60260-60472), mirroring the information in the GTF file- `species`: quoted species name- `prediction_program`: the predictor program that produced this gene structure (genewise, exonerate or blast)- `target`: the filename of the genome in fasta format--- ### GTF Files (`*.gtf`) These files contain the genomic intervals of selenoprotein predictions. The `gene_id` column follows the notation: profile_name.index.label (see above for explanation) Other data has been added:- `filter_status`: Wether selenoprofiles lineage has discarded the prediction or not. - `Annotation`: Annotation status of that selenoprofiles prediction.- `Orthology`: Subfamily classification by selenoprofiles orthology.- `Transcript_id_ens`: ID of annotations (NCBI or Ensembl) overlapping with this selenoprotein prediction. --- ### TSV Files (`*.tsv`) Tab-separated files saving selenoprotein IDs along with their Orthology label and Filtering status after the orthology and lineage utilities by selenoprofiles. --- ## Notes - The datasets are organized separately for NCBI and Ensembl to maintain clarity on source origin. - The filtering status in TSV files helps in downstream analyses to select high-confidence selenoprotein predictions. --- *For questions or clarifications, contact us at the email provided in the published manuscript*



