MD dataset
收藏资源简介:
Molecular Dynamics Dataset Description (for Molecular Property Prediction) This dataset comprises molecular dynamics (MD) simulation data, designed to support machine learning tasks involving molecular trajectories, especially molecular property prediction. The dataset covers two representative tasks: Toxicity Prediction and Protein-binder Identification All simulations were conducted using GROMACS, and the trajectories are stored in standard formats. Dataset and Task Types Dataset Number Format Task TOX21-MD 4490 .pdb Toxicity Prediction ADA17-MD 1176 .pdb + .xtc Protein-binder Identification EGFR-MD 1485 .pdb + .xtc Protein-binder Identification HIVPR-MD 994 .pdb + .xtc Protein-binder Identification Note: Due to Zenodo's upload size limit, the TOX21 dataset only includes multi-frame .pdb trajectory files. The other datasets (ADA17, EGFR, HIVPR) include both .pdb and .xtc files. File Format Description .xtc: Compressed trajectory files storing atomic coordinates over time steps, optimized for efficient storage and processing. .pdb: Multi-frame trajectory files for small molecules, suitable for 3D visualization and static structural analysis. This dataset supports the following tasks: Toxicity prediction based on structure or trajectory Molecular behavior modeling as time series Analysis of structural evolution and dynamic features Usage Notes You can extractatomic molecular topological structure, molecular conformational changes, some simulation parameter information, etc. Suitable for deep learning models such as CNNs, GNN and Transformer, etc. Download and Extraction Instructions Due to the large file sizes, all datasets are split into multiple zip parts. Please follow these guidelines for extraction: All zip parts (e.g., .z01, .z02, ...) must be placed in the same folder before extraction. Ensure at least 200 GB of free disk space for successful extraction. Citation If you use this dataset in your research, please cite our paper: title=' Dynamics-enhanced Molecular Property Prediction Guided by Deep Learning'



