遇见数据集

A Comprehensive Dataset for lncRNA-Disease-Gene Associations with CGR and FCGR Representations

收藏
Zenodo2025-05-01 更新2026-05-26 收录
官方服务:

资源简介:

lncRNA–Disease–Gene Association Dataset with CGR/FCGR Representations DescriptionThis dataset provides a standardized and curated resource for analyzing the associations between long non-coding RNAs (lncRNAs), diseases, and genes. It integrates data from three widely-used biological databases: - LncRNADisease- Comparative Toxicogenomics Database (CTD)- Ensembl It is designed to support applications in bioinformatics, systems biology, and machine learning by combining raw biological sequences, structured association matrices, and visual sequence encodings. --- Dataset Contents - FASTA – 2,594 lncRNA sequences in FASTA format - Associations/ – - 6,662 lncRNA–disease associations - 37,238 gene–disease associations - 256 diseases, 351 genes - Matrices/ – Binary and weighted association matrices for: - lncRNA–disease - gene–disease - Images/ – - CGR (Chaos Game Representation) images - FCGR (Frequency CGR) images - Scripts/ – - scraping.py: Automated extraction from source databases - preprocess.py: Cleaning, normalization, and formatting - generate_images.py: Sequence-to-image transformation (CGR/FCGR) --- Applications This dataset can be used for: - Predictive modeling of lncRNA-disease relationships - Network-based disease association analysis - Sequence-based and image-based machine learning models (e.g., CNNs, transformers) --- ## Reproducibility All scripts and data files are open-access and reproducible. To regenerate the data: python scraping.pypython preprocess.pypython generate_images.py Data Preparation Workflow https://zenodo.org/api/records/15316601/draft/files/data_pipeline_flowchart.png/contentThis diagram shows the six-step data processing pipeline, including extraction, preprocessing, matrix generation, and image transformation.

提供机构:
Zenodo
创建时间:
2025-05-01
二维码
社区交流群
二维码
科研交流群
商业服务