Natural language processing of gene descriptions for overrepresentation analysis with GeneTEA
收藏资源简介:
Data and models supporting "Natural language processing of gene descriptions for overrepresentation analysis with GeneTEA" (Boyle et al. 2025)File descriptions:GeneTEA.pkl, GeneTEA-yeast.pkl, PharmaTEA.pkl - pickled GeneTEA modelsFig2b/d.csv: Top terms in Figure 2 b/d.gProfiler_hsapiens_7-16-2025_4-35-07 PM__intersections.csv: g:GOSt results for Fig2b, downloaded from the g:Profiler site.gProfiler_hsapiens_7-16-2025_4-41-03 PM__intersections.csv: g:GOSt results for Fig2d, downloaded from the g:Profiler site.enrichr_sets_03_01_2025.csv: Enrichr database downloaded 3/1/2025, used for Figure 3 and S1.gene sets for connexin.gmt: Enrichr gene sets containing the term "connexin", downloaded from the Enrichr site.false_discoveries.csv: Benchmarking results for false discovery control in Figures 4 and S1.EF_hand_example.csv: Top terms and MedCPT scores for EF-hand example in Figure 4.[Hallmark or Experimentally Derived Queries]_scores.csv: Benchmarking results for [Hallmark or Experimentally Derived Queries] across joined top terms in Figures 4 and S2. The "joined_ranking" column corresponds to the MedCPT Relevance score across the top terms and "num_high_redundancy" contains the number of redundant term pairs.[Hallmark or Experimentally Derived Queries]_indiv.csv: Benchmarking results for [Hallmark or Experimentally Derived Queries] for each top term in Figures 4 and S2. The "indiv_ranking" column corresponds to the MedCPT Relevance score for a single term.Fig4_left/right: Examples of top terms shown in what is now Figure 5.gProfiler_hsapiens_3-13-2025_9-14-59 AM__intersections.csv: g:GOSt results for Fig4_left, downloaded from the g:Profiler site.gProfiler_hsapiens_2-11-2025_10-18-19 AM__intersections.csv: g:GOSt results for Fig4_right, downloaded from the g:Profiler site.



