Data from Barrel-aged SweetNet: Infusing a Graph Neural Network with Language Model Embeddings for Glycan Property Prediction
收藏资源简介:
This deposit contains the complete experimental data and associated files supporting the findings presented in the Bachelor's thesis titled "Barrel-aged SweetNet: Assessing the Utility of Infusing a Graph Neural Network with Language Model Embeddings for Glycan Property Prediction." Dataset Overview: The data originates from computational experiments designed to evaluate the impact of Glycan Language Model (GlyLM) embeddings on SweetNet, a Graph Neural Network (GNN), for glycan property prediction. Experiments were conducted across diverse biological tasks, including disease association (df_disease), species classification (df_species), and tissue sample prediction (df_tissue), utilizing data sourced from the glycowork library. Contents: The deposit includes data organized by experiment, covering: Performance Metrics: Detailed summary_*.csv files containing LRAP, NDCG, and Loss metrics for each run, along with aggregated statistics. Model State Dictionaries: Saved .pth files for trained and untrained models (baseline and infused configurations) allowing for full model inspection and replication. Test Set Data: Pickled (.pkl) files containing the held-out test sets (glycan sequences and labels) for each experiment, used for final, unbiased model evaluation. Experiment Settings: experiment_settings_*.yaml files, detailing all parameters and configurations used for each automated experiment batch, ensuring full reproducibility of experimental conditions. Supplementary Data: Includes additional analysis results such as comprehensive embedding performance comparisons across different GlyLM sources and detailed Euclidean distance analysis. Methodology & Key Findings: Experiments were executed using a custom-developed Hyperautomated Barrel-Batching System (HBBS), ensuring high standards of reproducibility. The study critically investigates the efficacy of GlyLM infusion, finding that contrary to the initial hypothesis, GlyLM-infused SweetNet models consistently exhibited reduced predictive performance compared to baseline models. Analysis suggests that embeddings may function more as simple glycoletter identifiers rather than complex semantic carriers under the tested conditions. Usage: This dataset facilitates the reproducibility of the computational experiments and further independent analysis. The associated project code repository (link available in the thesis and metadata) provides all necessary scripts and tools for data processing and model execution.



