遇见数据集

Supplementary Material – Master's Dissertation: Integration of Genomic and Toxicogenomic Data in the Investigation of Environmental Factors in Autism Spectrum Disorder

收藏
Zenodo2026-02-27 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the complete supplementary material associated with the Master’s dissertation entitled: “Integration of Genomic and Toxicogenomic Data in the Investigation of Environmental Factors in Autism Spectrum Disorder” Author: Bianca dos Reis SantosInstitution: Universidade Federal de Minas Gerais (UFMG)Year: 2026 Dataset Overview This dataset includes all supplementary tables, gene lists, SNP-level meta-analysis results, MAGMA gene-based findings, compound–gene interaction datasets derived from the Comparative Toxicogenomics Database (CTD), differential expression filtering results from GEO datasets, pathway enrichment analyses, network-based prioritization metrics, and regulatory element mapping. The material is organized according to the dissertation structure, with clear separation between the two main analytical chapters: Chapter 2 – Integrated Prioritization Pipeline: From GWAS to Selective Compounds This chapter presents a multi-layer integrative strategy to prioritize chemical compounds with selective interaction patterns toward autism-related genes. The supplementary material includes: GWAS Meta-analysis Outputs: Summary statistics from 18 GWAS datasets, LD Score regression results, and FUMA-annotated gene lists. Compound–Gene Interaction Datasets: Full and filtered interaction sets from the CTD, including hypergeometric enrichment tests for both GWAS-derived and SFARI-derived gene lists. Machine Learning Filtering: Random Forest classification outputs (ECFP4 fingerprints + physicochemical properties), including probability scores (RF_proba) and final prioritization categories. Transcriptomic Filtering: Differential expression results from four GEO brain tissue datasets, with gene priority classification (Gold/Silver/Copper). Network Topology Metrics: Centrality measures (degree, betweenness, closeness) from STRING-derived PPI networks. Pathway Enrichment: Full and semantically reduced GO/KEGG/Reactome terms for gene sets at multiple stages. Final Integrated Tables: Compound–gene–pathway master table, ranked compounds by Compound Impact Score (CIS), prioritized pathways, and final gene lists with integrated centrality metrics. Chapter 3 – Regulatory Architecture and Gene–Compound Integration This chapter expands the analysis by mapping genomic regulatory elements associated with prioritized genes and identifying potential cis-regulatory disruption mechanisms. The supplementary material includes: Regulatory Element Mapping: Genome-wide associations between prioritized genes and brain-active enhancers (cerebellar and cortical) and long non-coding RNAs (lncRNAs) within ±100 kb windows. SNP–Regulatory Overlap: Identification of GWAS-derived risk variants overlapping enhancers and lncRNA loci, with precise genomic coordinates (hg38). Hotspot Gene Identification: Ranking of genes by regulatory density (total enhancer + lncRNA counts), defining regulatory hotspots (≥90th percentile). Integrated Gene–Compound–Regulation Table: Final cross-sectional output linking selective compounds (from Chapter 2) to hotspot genes and their associated regulatory elements, including SNP overlap counts and literature references. Supplementary Figures and Additional Data Figures: QQ plots, Manhattan plots, PCA of enriched pathways. Methodological Details: Computational protocols (e.g., bedtools commands for genomic intersections). File Naming and Traceability All files are named according to their reference in the dissertation (e.g., Table S1.xlsx, Table S2.csv, Figure S1.png) to ensure direct and unambiguous traceability between the thesis text and the deposited material. These files are provided to ensure full transparency, reproducibility, and open access to all computational analyses performed in this work. Code Availability-----------------All custom scripts and pipelines developed for this work are available in the 'Projeto_Pipeline_Integrativo' directory within this archive. They include: - GWAS preprocessing and meta‑analysis (Python/R)- LD Score regression wrapper- FUMA post‑processing- CTD chemical‑gene interaction analysis and hypergeometric tests- Random Forest classifier for compound prioritization- Transcriptomic filtering (DEGs from GEO datasets)- Network analysis (centrality, clustering, STRING integration)- Pathway enrichment and semantic reduction (g:Profiler, REVIGO)- Regulatory element mapping (enhancers, lncRNAs) Each script contains detailed comments and there are txt files explaining requirements for each step. Refer to the dissertation text for detailed methodological descriptions and file naming conventions.

提供机构:
Zenodo
创建时间:
2026-02-16
二维码
社区交流群
二维码
科研交流群
商业服务