High quality protein residues: Top2018 mainchain-filtered residues - mmCIF
收藏资源简介:
Introduction<br> --------------------------------------------------------------------------------<br> This directory contains files from the Top2018 dataset by the Richardson Lab at Duke University. This version contains files in mmCIF format. These are high-quality residues from high-quality, low redundancy protein chains in the PDB. This dataset is quality-filtered on mainchain atoms. The accompanying publication is:<br> Williams, C. J., Richardson, D. C., & Richardson, J. S. (2021). The importance of residue-level filtering, and the Top2018 best-parts dataset of high‐quality protein residues. Protein Science. http://doi.org/10.1002/pro.4239 Usage recommendations<br> --------------------------------------------------------------------------------<br> Protein residues that fail the filtering criteria described below have been removed from the files. As a result, these files can be considered pre-filtered and will return only results for residues of good model quality with supporting experimental data. As long as the question concerns mainchain protein atoms, these files should be usable as is. There is a separate version that has been filtered on all atoms that is suitable for sidechains. The Top2018 contains several different levels of homology clustering (30%, 50%, 70%, 90%) to ensure nonredundant datasets. The 70% homology level is a reliable default. These chains are listed in top2018_chains_hom70_mcfilter_60pct_complete.txt and found in top2018_cifs_mc_filtered_hom70.tar.gz Files are organized in subdirectories based on the first two letters of their PDB ids. The included python script sample_file_loop.py may aid in accessing the directory structure. Files already contain hydrogens added by Reduce. NQH flips have been performed to ensure that these are the best versions of these structures. top2018_metadata_mc_filtered.csv contains information on release date, resolution, and validation scores for each file. top2018_passrates_mc_filtered.csv contains information on how many protein residues from the original chain passed the quality filters. <br> Homology sets:<br> --------------------------------------------------------------------------------<br> Using sequence homology clusters provided by the RCSB PDB, for each homology cluster, the best chain was selected for inclusion in the dataset. This ensures minimal sequence/structural redundancy. The Top2018 is available at several different levels of homology clustering, which may be appropriate to different uses. Lists of the included chains at each homology level are included in this distribution. Lower homology numbers mean less redundancy, but fewer total chains in the dataset. For general use, ***we recommend the 70% homology set*** as a good balance between inclusivity and variety. This list is given in the file top2018_chains_hom70_mcfilter_60pct_complete.txt <br> Usage caveats:<br> --------------------------------------------------------------------------------<br> These files are incomplete. They are single chains from structures that may have had multiple chains. Residues that fail the filtering criteria have been removed. Programs with strong requirements for completeness or uninterrupted chains should be used with care. Chain completeness and fragmentation statistics are available in top2018_passrates_mc_filted.csv and as _top2018.percent_passrate in the .cif file. All ligands and waters associated with the chain have been preserved without filtering. Robust ligand filtering is beyond the scope of this dataset. Trust the ligands at your own discretion. Sidechain atoms beyond CB have not been considered in the filtering. However, all sidechains have been included for residues that passed the mainchain filters. DO NOT use this set of files for serious questions involving sidechains. See our all-atom filtered dataset instead. <br> Filtering criteria: Chain level<br> --------------------------------------------------------------------------------<br> Chain is protein<br> Released on or before Dec 31, 2018<br> Resolution < 2.0<br> MolProbity Score < 2.0<br> <3% residues have cbeta deviations<br> <2% residues have covalent bond length outliers<br> <2% residues have covalent bond geometry outliers Using sequence homology clusters provided by the RCSB PDB, for each homology cluster, the chain with the best (lowest) average of Resolution and MolProbity Score was selected. <br> Filtering criteria: Residue level<br> --------------------------------------------------------------------------------<br> Even excellent structures usually contain some poorly-resolved regions. Residue-level filtering helps avoid including these regions in otherwise high-quality data Mainchain atoms are defined as N, CA, C, O, CB.<br> Note that CB is included, since its ideal position is defined by the other mainchain atoms. All mainchain atoms in a residue:<br> Bfactor <= 40<br> Real-space correlation coefficient (rscc) >= 0.7<br> 2Fo-Fc map value >= 1.2 Additionally, residues are not allowed to have:<br> Covalent geometry outliers<br> Steric overlaps or "clashes", as per Probe<br> Alternate conformations <br> Chain Completeness criteria<br> --------------------------------------------------------------------------------<br> Chains which lost >40% of their residues during filtering were dropped from this dataset. All chains present here are at least 60% complete. <br> Filtering doumentation<br> --------------------------------------------------------------------------------<br> Each file documents its pruned residues and included segments in a cif data block named data_top2018_dataset. This block can be found at the end of the file. In the _top2018_deleted_residue loop, causes of pruning are documented. If a residues was removed due to failing the B-factor filter, a "b" will appear in the appropriate column. Otherwise, a "." will appear. Other filtering criteria are treated similarly with the following codes:<br> b - B-factor<br> c - RSCC<br> m - map value<br> g - geometry outliers<br> o - steric overlaps<br> a - alternate conformations Version history<br> --------------------------------------------------------------------------------<br> Version 0.9<br> Initial upload to establish DOI Version 1.0<br> Initial complete upload
### 简介 -------------------------------------------------------------------------------- 本目录包含杜克大学理查森实验室(Richardson Lab)发布的Top2018数据集相关文件。本版本提供mmCIF格式的文件,数据源自蛋白质数据库(Protein Data Bank,PDB)中高质量、低冗余的蛋白质链,提取得到高品质残基。本数据集针对主链原子进行了质量过滤。配套发表的研究论文为:Williams C J, Richardson D C, Richardson J S. 残基级别过滤的重要性与Top2018高品质蛋白质残基最优子集数据集[J]. 蛋白质科学, 2021. https://doi.org/10.1002/pro.4239 ### 使用建议 -------------------------------------------------------------------------------- 所有未通过下述过滤标准的蛋白质残基均已从文件中移除,因此本数据集已完成预过滤,仅保留具备优质模型质量且有实验数据支撑的残基。若研究仅涉及蛋白质主链原子,可直接使用本数据集文件。另有针对所有原子完成过滤的版本,适用于侧链相关研究。 Top2018数据集提供了多种同源聚类层级(30%、50%、70%、90%)以保障数据集的非冗余性,其中70%同源聚类为可靠的默认选择。符合该层级的蛋白质链信息存储于top2018_chains_hom70_mcfilter_60pct_complete.txt,对应压缩包为top2018_cifs_mc_filtered_hom70.tar.gz。 文件按照PDB ID的前两个字母分至不同子目录中。附带的Python脚本sample_file_loop.py可辅助遍历目录结构。所有文件均已通过Reduce工具添加氢原子,并完成了NQH翻转操作,以确保结构为最优版本。 top2018_metadata_mc_filtered.csv文件存储了每个文件的发布日期、分辨率及验证评分信息;top2018_passrates_mc_filtered.csv则记录了原始蛋白质链中通过质量过滤的残基占比情况。 ### 同源聚类集 -------------------------------------------------------------------------------- 本数据集采用RCSB PDB提供的序列同源聚类结果,针对每个同源聚类簇选取最优链纳入数据集,以最大限度降低序列与结构冗余度。Top2018数据集提供多种同源聚类层级,可适配不同研究需求,各聚类层级对应的入选蛋白质链列表均随本数据包一同发布。同源阈值越低,数据集冗余度越低,但总链数也越少。通用研究场景下,**我们推荐选用70%同源聚类集**,其在数据包容性与多样性间取得了良好平衡,对应链列表存储于top2018_chains_hom70_mcfilter_60pct_complete.txt文件中。 ### 使用注意事项 -------------------------------------------------------------------------------- 本数据集文件并非完整结构,仅提取自原结构中的单条蛋白质链(原结构可能包含多条链),且已移除未通过过滤标准的残基。若研究对链的完整性或无中断序列有严格要求,请谨慎使用本数据集。链完整性与片段化统计信息可从top2018_passrates_mc_filtered.csv文件以及.cif文件中的_top2018.percent_passrate字段获取。与目标链相关的所有配体与水分子均未经过滤直接保留,本数据集未涵盖配体的严格过滤流程,配体信息的可靠性需自行判断。过滤过程未考虑CB原子以外的侧链原子,但所有通过主链过滤的残基的侧链均被保留。**请勿将本数据集用于涉及侧链的严谨研究,请改用全原子过滤版本的数据集**。 ### 链级别过滤标准 -------------------------------------------------------------------------------- 1. 目标链为蛋白质链 2. 发布日期不晚于2018年12月31日 3. 分辨率<2.0 Å 4. MolProbity评分<2.0 5. 存在CB原子偏差的残基占比<3% 6. 存在共价键长度异常的残基占比<2% 7. 存在共价键几何异常的残基占比<2% 此外,针对每个同源聚类簇,选取分辨率与MolProbity评分的平均值最低的链纳入数据集。 ### 残基级别过滤标准 -------------------------------------------------------------------------------- 即便优质的蛋白质结构也往往存在部分分辨率较差的区域,残基级别过滤可避免将这些区域纳入整体高质量数据中。主链原子定义为N、CA、C、O及CB原子(CB原子的理想位置由其他主链原子确定,故纳入主链范畴)。 单条残基的所有主链原子需满足以下条件: 1. B因子≤40 2. 实空间相关系数(Real-space correlation coefficient, RSCC)≥0.7 3. 2Fo-Fc映射值≥1.2 此外,残基不得存在以下问题: 1. 共价几何异常 2. 由Probe工具检测到的空间重叠或“冲突”(clashes) 3. 存在交替构象 ### 链完整性标准 -------------------------------------------------------------------------------- 过滤过程中丢失残基占比>40%的链已被从数据集中移除,本数据集包含的所有链的完整性至少为60%。 ### 过滤文档 -------------------------------------------------------------------------------- 每个文件均在名为data_top2018_dataset的CIF数据块中记录了被移除的残基与保留的片段,该数据块位于文件末尾。在_top2018_deleted_residue循环区块中,记录了残基被移除的原因:若残基因B因子过滤被移除,则对应列标注为“b”,否则标注为“.”。其他过滤标准的标注代码如下: - b:B因子过滤 - c:RSCC过滤 - m:映射值过滤 - g:几何异常过滤 - o:空间重叠过滤 - a:交替构象过滤 ### 版本历史 -------------------------------------------------------------------------------- - 版本0.9:首次上传以获取DOI标识 - 版本1.0:首次完整上传



