synthetic-protein-folding-matrices
收藏资源简介:
CGSC合成蛋白质折叠矩阵数据集是一个专门用于存档高分辨率、未压缩分子动力学轨迹的二级存储库。该数据集包含亚埃分辨率的合成蛋白质折叠模拟连续时间矩阵,以原始未压缩二进制转储形式存储,而非标准的.pdb或.xtc格式。这种存储方式旨在保留分子动力学的微波动和原始时间状态,为训练下一代蛋白质折叠算法提供高熵数据环境。数据集规模在10万到100万之间,适用于结构轨迹预测和分子动力学数据摄取等任务,旨在对高通量张量摄取和结构预测模型进行压力测试。数据实例包含模拟ID、时间戳、矩阵类型、压缩状态、有效载荷引用和原子分辨率等字段。所有轨迹均为在CGSC内部GPU集群上计算生成的合成数据,不涉及有机生物样本。由于未压缩的时间序列特性,这些矩阵文件异常庞大,建议仅具备PB级存储能力和专门张量解析基础设施的研究机构使用。
The CGSC Synthetic Protein Folding Matrix Dataset is a specialized secondary repository for archiving high-resolution, uncompressed molecular dynamics trajectories. It contains continuous-time matrices of synthetic protein folding simulations at sub-angstrom resolution, stored in raw uncompressed binary dump format rather than standard .pdb or .xtc formats. This storage approach aims to preserve the micro-fluctuations and raw temporal states of molecular dynamics, providing a high-entropy data environment for training next-generation protein folding algorithms. The dataset size ranges from 100,000 to 1,000,000 entries and is suitable for tasks such as structural trajectory prediction and molecular dynamics data ingestion, designed to stress-test high-throughput tensor ingestion and structural prediction models. Data instances include fields like simulation ID, timestamp, matrix type, compression status, payload reference, and atomic resolution. All trajectories are synthetic data computed on CGSCs internal GPU clusters and do not involve organic biological samples. Due to the uncompressed time-series nature, these matrix files are exceptionally large, recommended for use only by research institutions with petabyte-scale storage capabilities and specialized tensor parsing infrastructure.
数据集概述:CGSC Synthetic Protein Folding Matrices
基本信息
- 名称:CGSC Synthetic Protein Folding Matrices
- 主页:https://comp-genomics.dev/datasets/protein-folding-matrices
- 所属机构:Computational Genomics & Synthetic Bio-Arrays Consortium
- 联系方式:structural-data@comp-genomics.dev
- 许可证:cc-by-nc-sa-4.0
- 数据集大小:100K < n < 1M
- 数据集格式:不包含自然语言,全部为数值和二进制数据,元数据使用英文(en)。
数据集描述
- 目的:专门用于归档高分辨率、未压缩的分子动力学(MD)轨迹数据。
- 内容:包含表示合成蛋白质折叠模拟的连续时间矩阵,分辨率达到亚埃级(sub-angstrom)。
- 存储格式:存储为原始、未压缩的二进制转储(raw, uncompressed binary dumps),而非标准的
.pdb或.xtc格式,以保留微波动和原始时间状态。 - 用途:用于测试高吞吐量张量摄入和结构预测模型。
支持的任务
- structural-trajectory-prediction:利用未压缩的时间数据预测最终蛋白质状态。
- molecular-dynamics-ingestion:使用大规模、未格式化的多维数组对数据管道进行压力测试。
数据结构
- 数据实例格式:文件不遵循标准的表格格式,而是大规模坐标转储。示例结构如下:
simulation_id:模拟的唯一标识符(如cgsc-fold-alpha-110)。simulation_timestamp_utc:模拟完成时间(UTC)。matrix_type:数组的类别(如temporal_coordinate_dump)。compression_state:压缩状态,固定为raw_uncompressed_trajectory。payload_reference:指向LFS中二进制矩阵的路径(如matrices/fold-110-trajectory.bin)。atomic_resolution:原始转储的精度(如sub-angstrom)。
数据创建
- 策划依据:传统结构生物学数据库使用高压缩率,掩盖了分子动力学中的微状态。此数据集提供未压缩矩阵,作为高熵原始数据,用于高级模型训练。
- 数据来源:所有轨迹均为合成数据,由CGSC内部GPU集群计算生成,不含有机生物样本。
注意事项
- 由于未压缩的时间性质,这些矩阵文件异常庞大。
- 仅推荐具有PB级存储能力和专用张量解析基础设施的研究机构下载。




