Ontologizer Dataset
收藏资源简介:
Simulated gene sets for Gene Ontology overrepresentation analysis, for five commonly studied organisms. The data were generated to accompany the Ontologizer 3 application. The top-level archive contains the Gene Ontology (go-basic.json) and the five organism-specific GO annotation files (*.gaf.gz). Each organism's subfolder contains three files: population_genes.txt The population gene set, comprising all protein-coding genes of the organism, one gene symbol per line. study_genges.txt The simulated study gene set drawn from the population set. The study set is constructed by sampling Gene Ontology terms and adding a fraction of each term's annotated genes and a number of unrelated noise genes. solution.tsv A tab-separated ground-truth file with two columns: the GO term ID (or the label `Noise` for unrelated genes), and a comma-separated list of the genes drawn from that term and added to the study set. These datasets are intended for exploring the methods provided by Ontologizer on data with a known causal structure.



