遇见数据集

A Consensus of In-silico Sequence-based Modeling Techniques for Compound-Viral Protein Activity Prediction for SARS-COV-2

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Here we provide the datasets used for training and testing of the end-to-end supervised deep learning models as well as the datasets used with vector representations of compounds and proteins and passed to supervised state-of-the-art machine learning models (XGBoost, RF, SVM). We also provide the full list of viral proteins with their sequences used for the protein autoencoder along with the list of SMILES representations of compounds used for the compound autoencoder. Furthermore, we provide pickle files of data obtained from NCBI assay and compound-viral protein interactions downloaded through ChEMBL. The compound-viral protein interactions after filtering from both NCBI and ChEMBL. The list of compounds tested against the three main proteases of coronavirus and the three main proteases of SARS-COV-2 as a fasta file. All the test files associated with SARS-COV-2 viral proteins for end-to-end deep learning models as well as vector representation based supervised machine learning models.

本研究提供了用于端到端监督深度学习模型训练与测试的数据集,以及结合化合物与蛋白质向量表征、输入至当前最优监督机器学习模型(XGBoost、RF、SVM)的数据集。此外,本研究还提供了用于蛋白质自编码器(protein autoencoder)的完整病毒蛋白质序列列表,以及用于化合物自编码器的化合物SMILES表征列表。进一步地,本研究提供了从NCBI实验获取的数据以及通过ChEMBL下载的化合物-病毒蛋白质相互作用数据的pickle文件,同时包含经NCBI与ChEMBL双重筛选后的化合物-病毒蛋白质相互作用数据集。此外还提供以FASTA格式存储的、针对冠状病毒三类主要蛋白酶与SARS-CoV-2三类主要蛋白酶进行测试的化合物列表,以及所有针对SARS-CoV-2病毒蛋白质、适配端到端深度学习模型与基于向量表征的监督机器学习模型的测试文件。

创建时间:
2020-11-03
二维码
社区交流群
二维码
科研交流群
商业服务