遇见数据集

Supplemental datasets for: 'Pan-microalgal dark proteome classification via interpretable deep learning with synthetic chimeras'

收藏
Zenodo2025-05-31 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains supplementary datasets for 'Pan-microalgal dark proteome classification via interpretable deep learning with synthetic chimeras'. The preprint is available at https://doi.org/10.48550/arXiv.2411.06798. Data S1 | Data S1 | Neural network training sequences (4GB). Dataset comprising DNA sequences processed in two formats: with terminal information (TI-inclusive) and terminal information-free (TI-free). Data S2 | Training and inference scripts. The LA4SR framework integrated several open-source software packages and models. We employed LORA (Low-Rank Adaptation (https://github.com/microsoft/LoRA)), PEFT (Parameter-Efficient Fine-Tuning (https://github.com/huggingface/peft), and QLORA (Quantized Low-Rank Adaptation (https://github.com/artidoro/qlora) for parameter-efficient post-training and used Mamba (https://github.com/state-spaces/mamba) as an alternative to transformer-based architectures. The Hugging Face Transformers library (https://github.com/huggingface/transformers) facilitated implementation, pretraining, and post-training of the open-source models. Training and inference were carried out on the High Performance Computing resources at New York University Abu Dhabi (Jubail HPC cluster), with jobs going to NVIDIA (Santa Clara, CA, USA) V100, A100, or H100 nodes. Data S3 | Interpretability scripts, including DeepMotifMinerPro. 10.5281/zenodo.13920001. Includes scripts for the implementation of the custom explainer programs presented with this work, including Captum-, DeepLift, and SHAP-based approaches (Data S3) to explain how different amino acid residues and their patterns and positions affect model decisions. Data S4 | Real-world sequencing data. To validate our approach and address real-world challenges, we applied LA4SR models to new data from new clean and contaminated isogenic algae cultures. We cultured and sequenced ten separate isogenic colonies of Chlamydomonas reinhardtii CC-1883. Of these, nine were sequenced with Illumina 150 bp paired-end short reads and one with Pacific Biosciences (PacBio, Menlo Park, CA, USA) HiFi reads and DoveTail (Sydney, Australia) Hi-C to generate a complete, axenic reference assembly (Fig. S3; Data S4).

本仓库为《基于合成嵌合体可解释深度学习的泛微藻暗蛋白质组分类》研究提供补充数据集。该预印本可通过https://doi.org/10.48550/arXiv.2411.06798获取。 数据集S1 | 神经网络训练序列(4GB)。本数据集包含经两种格式处理的DNA序列:带末端信息(TI-inclusive)与无末端信息(TI-free)。 数据集S2 | 训练与推理脚本。LA4SR框架集成了多款开源软件包与模型。本研究采用LORA(Low-Rank Adaptation,低秩适配)、PEFT(Parameter-Efficient Fine-Tuning,参数高效微调)以及QLORA(Quantized Low-Rank Adaptation,量化低秩适配)实现参数高效的后训练,并使用Mamba作为基于Transformer架构的替代方案。依托Hugging Face Transformers库(https://github.com/huggingface/transformers)完成开源模型的实现、预训练与后训练流程。训练与推理工作均在纽约大学阿布扎比分校的高性能计算资源(Jubail HPC集群)上执行,作业调度至NVIDIA(美国加利福尼亚州圣克拉拉)V100、A100或H100计算节点。 数据集S3 | 可解释性脚本,包含DeepMotifMinerPro。相关资源DOI为10.5281/zenodo.13920001。本数据集包含实现本研究自定义解释程序的脚本,涵盖基于Captum、DeepLift以及SHAP的分析方法,用于阐释不同氨基酸残基及其模式与位置如何影响模型决策。 数据集S4 | 真实世界测序数据。为验证本研究方法并应对真实场景挑战,我们将LA4SR模型应用于来自清洁与污染的同基因藻类培养物的新数据。本研究培养并测序了10株独立的莱茵衣藻(Chlamydomonas reinhardtii)CC-1883同基因菌落。其中9株采用Illumina 150 bp双端短读长测序,1株采用太平洋生物科学公司(PacBio,美国加利福尼亚州门洛帕克)HiFi测序与澳大利亚悉尼DoveTail公司的Hi-C技术,以生成完整的无菌参考组装序列(图S3;数据集S4)。

提供机构:
Zenodo
创建时间:
2025-02-13
二维码
社区交流群
二维码
科研交流群
商业服务