Dataset of AlphaFold 3 and RoseTTAFold2NA Structures for Analysing the Capabilities of These Programmes in Predicting the Spacing and Orientation Preferences of Two Transcription Factors
收藏资源简介:
Description:This dataset contains a collection of predicted TF‑TF–DNA complex structures generated using two state‑of‑the‑art artificial intelligence programmes: RoseTTAFold2NA and AlphaFold 3. The structures were predicted as part of a study aimed at evaluating the capability of these models to capture the spacing and orientation preferences in TF‑TF–DNA interactions, as derived from CAP‑SELEX experimental data. This repository is accompanied by an associated Zenodo repository [https://doi.org/10.5281/zenodo.14846538] containing the code used for analysing these data. The code repository was automatically copied by Zenodo from this GitHub repository Background and Rationale:Transcription factors bind to specific DNA sequences to regulate gene expression. Experimental approaches such as CAP‑SELEX have enabled the characterisation of binding preferences by providing enriched k‑mer motifs and information on preferred spatial arrangements between TF binding sites. In this study, 119 position weight matrix (PWM) models, obtained from CAP‑SELEX experiments, were used to define the preferred DNA sequences for pairs of transcription factors (e.g. PAX2 and ELK3). Dataset Content:A total of 1,620 PDB files were generated for each program, resulting in 3,240 predicted structures overall. The sequences used as input for AlphaFold 3 were identical to those used for RoseTTAFold2NA. For each PWM model, multiple competitive DNA sequences were generated to probe different spatial and orientational configurations. Methodology:The prediction pipeline involved the following key steps: Derivation of Input Sequences: Based on CAP‑SELEX data, two key k‑mers were identified for each TF. Multiple DNA sequences were constructed by varying the spacing (e.g. ±1 or ±2 nucleotides) and the orientation (e.g. swapping positions or using reverse complements) between the two k‑mers. Structure Prediction: The generated sequences were used as input for RoseTTAFold2NA and AlphaFold 3, resulting in a set of predicted TF‑TF–DNA complex structures. Analysis: The predicted structures were analysed by comparing the number of contacts in the region corresponding to the preferred DNA versus the competitive DNA. A contact was defined as a pair of atoms (one from an amino acid and one from a nucleic acid) whose minimum distance is less than the threshold of 0.45 nm. By linking experimental CAP‑SELEX data with high‑accuracy structural predictions, this resource facilitates a deeper understanding of the molecular mechanisms underlying transcriptional regulation and provides a valuable resource for improving existing methods as well as serving as a dataset for further methodological development.
数据集说明:本数据集收录了利用两款当前顶尖人工智能程序——RoseTTAFold2NA与AlphaFold 3——预测得到的转录因子-转录因子-DNA(TF-TF-DNA)复合物结构集合。本数据集的结构预测工作隶属于一项研究,该研究旨在评估上述模型捕捉转录因子-转录因子-DNA相互作用中间隔与取向偏好的能力,相关实验数据源自CAP-SELEX实验。本仓库附带一个关联的Zenodo知识库[https://doi.org/10.5281/zenodo.14846538],其中包含用于分析本数据集的代码。该代码仓库由Zenodo从本GitHub仓库自动复刻而来。 背景与研究依据:转录因子通过结合特定DNA序列调控基因表达。诸如CAP-SELEX这类实验方法,可通过富集得到的k聚体(k-mer)基序以及转录因子结合位点间优选空间排布信息,实现对结合偏好性的表征。本研究采用了从CAP-SELEX实验中获取的119个位置权重矩阵(Position Weight Matrix, PWM)模型,用于定义成对转录因子(例如PAX2与ELK3)的优选DNA序列。 数据集内容:针对每一款人工智能程序,均生成了1620个PDB格式文件,总计得到3240个预测复合物结构。用于AlphaFold 3的输入序列与用于RoseTTAFold2NA的输入序列完全一致。针对每个位置权重矩阵模型,均生成了多条竞争性DNA序列,以探究不同的空间与取向构型。 研究方法:预测流程包含以下关键步骤: 输入序列的推导:基于CAP-SELEX实验数据,为每个转录因子识别出两个关键k聚体。通过调整两个k聚体之间的间隔(例如±1或±2个核苷酸)与取向(例如交换位置或使用反向互补序列),构建得到多条DNA序列。 结构预测:将构建得到的序列作为输入,分别送入RoseTTAFold2NA与AlphaFold 3中进行预测,最终得到一批转录因子-转录因子-DNA复合物预测结构。 数据分析:通过对比偏好DNA区域与竞争性DNA区域的原子接触数,对预测得到的复合物结构进行分析。此处的原子接触定义为:一个来自氨基酸、一个来自核酸的两个原子之间的最小距离小于0.45纳米的原子对。 通过将CAP-SELEX实验数据与高精度结构预测相结合,本资源有助于更深入地理解转录调控背后的分子机制,可为现有方法的优化提供有价值的参考,同时也可作为进一步方法学开发的数据集使用。



