遇见数据集

Supplementary Data and Models for: "Evaluation of a structure-based method for ab initio gene detection using deep learning"

收藏
Zenodo2026-03-19 更新2026-05-26 收录
官方服务:

资源简介:

Overview: This record contains a subset of the data and model assets for the research project described in the preprint "Evaluation of a structure-based method for ab initio gene detection using deep learning" by Hummel and Estrada (DOI: 10.64898/2025.12.19.694709). These files are supplementary to the primary research compendium on GitHub, which includes all source code, notebooks, and smaller data files. Primary Repository: To use these files, please refer to the documentation in the main GitHub repository for this research project (https://github.com/GreenMtn-bioinfo/RBIF120-HEX-finder-training). File Contents: The assets are organized into five categorical ZIP files to maintain the project's internal directory structure on GitHub while adhering to file size constraints: 1_Exon_Annotation.zip: GFF files containing RefSeq MANE Select exon features extracted from the annotation file for GRCh38.p14 (available on the NCBI website). The reference features of interest have been separated by strand to facilitate subsequent steps in preparing the training and testing data. 2_Selected_Coords_Seqs.zip: Files containing coordinates for each class (and sub-class) of sequence sampled to create the training and testing sets. These were used with samtools to retrieve the sequences for which physicochemical/structural profiles were calculated. This ZIP includes only the coordinate files that were >25 MB. The rest are hosted in the primary research compendium on GitHub (see above). 3_Physicochemical_Profiles.zip: NPY and JSON files that track important metadata about the physicochemical profiles, such as the unique IDs assigned to each profile (based on their source coordinate), the training/testing partition, class/sub-class labels, etc. Importantly, the file that contains the actual physicochemical profiles (named "all_profiles.npy") is not included in this ZIP. That file is ~90 GB, so it exceeds the Zenodo and GitHub file size limits. Please see the documentation for the main research compendium on how to recreate that file yourself using the provided code and data. 4_Sliding_Window.zip: NPY files >25 MB containing the intermediate pipeline data for each of the three sequence sets used in the "sliding window" evaluation of the models and post-processing pipeline, which was discussed in the preprint. The intermediate data in this ZIP include the structural profiles for the sequences, nucleotide-level probabilities from each model, and the filtered nucleotide-level predictions, which were further filtered into exon-level predictions by the post-processing pipeline. ChemEXIN_modified.zip: The original, unaltered model weights from the ChemEXIN repository on GitHub, which was the outcome of the work described in "Exon–intron boundary detection made easy by physicochemical properties of DNA" by Sharma et al. in 2025 (DOI: 10.1039/d4mo00241e). These were included here for stability of the primary research repository on GitHub, which contains a clone of the ChemEXIN repository that depends on these weights. Please see the Attribution & Licensing section below for more details. NOTE: The ZIP files only occupy a combined disk space of ~3.9 GB; however, once decompressed, the files will occupy a total of ~11 GB. Also, if you plan to run the code on GitHub to recreate the structural profiles for the training/testing set, you will need an additional ~90 GB of available disk space (for a total of over ~102 GB). Attribution & Licensing: Original Work: This project was inspired by work, ideas, and findings originally published by Sharma et al. and Mishra et al. (please see "References" and "Related Works" for this upload). While most of the data and code represent my own work on implementing the same underlying idea, a benchmarking portion of the research described in the preprint utilized the ChemEXIN tool by Sharma et al., which was the outcome of their latest published work (DOI: 10.1039/d4mo00241e). Third-Party Assets: The file ChemEXIN_modified.zip contains model weights from the aforementioned tool. These are included here strictly for stability and reproducibility, ensuring the research reflected in the primary repository for this project can be replicated exactly as conducted. Ownership: I do not claim ownership of these specific third-party weights; they are redistributed under their original GPL v3 License. Modifications: The weights included here are unchanged from those hosted on GitHub at the time this research was conducted. That said, some of the ChemEXIN code was modified to enable a benchmarking evaluation on a common sequence set. Detailed documentation of the changes made to external code can be found in the MODIFICATIONS.md file within the ChemEXIN clone in the primary GitHub repository for this project (see the repository documentation for more details).

提供机构:
Zenodo
创建时间:
2026-03-19
二维码
社区交流群
二维码
科研交流群
商业服务