RNA-SyntHub
收藏资源简介:
The structural diversity of RNA plays a pivotal role in molecular medicine, RNA-targeted drug discovery, and synthetic biology, where engineered RNA molecules are increasingly applied in therapeutic, regulatory, and nanotechnology contexts. The rapid development of deep learning methods has further increased the demand for large and diverse structural datasets; however, while protein datasets are abundant, high-quality RNA data remain scarce. Synthetic RNA structure datasets provide a means to alleviate this limitation by enabling the large-scale exploration of conformational space; however, their practical use requires careful curation and validation to ensure reliability. To address this challenge, we present a pipeline specifically designed for the systematic construction of curated synthetic RNA datasets. The workflow makes use of data provided by multiple structure generation approaches and integrates them with a meta-scoring function that discriminates low-quality models and extracts high-confidence candidates. As a proof of concept, we applied this method to a dataset of 447,402 synthetic RNA structures, obtaining a curated set of approximately 16,000 models, further complemented with predictions from Boltz-1 and RNAComposer. This study introduces both a methodological framework for making synthetic RNA datasets usable in computational research and a concrete dataset that can directly support deep learning applications and the advancement of RNA design in synthetic biology and nanotechnology.



