Assessing the many aspects of protein-coding gene annotation quality with OMArk
收藏资源简介:
Dataset associated to the OMArk paper. Contain four archives: SuppTables The Supplementary Table files referred to in the paper OMAmerDB: The OMAmer database constructed using the whole dataset of the OMA database (December 2021 Release) and used in the paper. An OMAmer database is necessary to run OMArk. Simulation:<br> Proteomes with artificially introduced errors, contaminants or depleted completeness, used to assess OMArk's performance. The archive contains the generated proteomes (Simulated_Data*) and their OMArk quality assessments (OMArk_Results). They also contains the OMAmer results (OMAmer_Placements) that were used to run OMArk and BUSCO completeness assessments (BUSCO_Results). *Note that for storage efficiency, only the non-redundant part of the data (added errors, added contamination, random fraction of proteomes) are stored there. The full modified proteome can be regenerated from these data and the source proteomes. Reference Proteomes: The UniProt Reference Proteomes (Proteomes) (2021_04) and their proteome quality assesment results according to OMArk. The archive also contains the OMAmer results (OMAmerResults) that were used to run OMArk (OMArk_Results), and BUSCO completeness assesments (BUSCO_Results).
本数据集关联OMArk论文,包含四个归档文件: 1. SuppTables:论文中提及的补充表格文件。 2. OMAmerDB:基于OMA数据库(OMA Database)2021年12月版全数据集构建的OMAmer数据库(OMAmer Database),为本论文所用。运行OMArk需依托该数据库。 3. 模拟数据集:包含人为引入错误、污染物或完整性缺损的蛋白质组,用于评估OMArk的性能。该归档包含生成的蛋白质组(Simulated_Data*)及其OMArk质量评估结果(OMArk_Results),同时还包含运行OMArk所需的OMAmer分析结果(OMAmer_Placements)以及BUSCO完整性评估结果(BUSCO)。*注:为优化存储效率,本归档仅留存数据的非冗余部分(包括引入的错误、新增的污染物、蛋白质组的随机子集)。完整的修改后蛋白质组可通过这些数据与源蛋白质组重新生成。 4. 参考蛋白质组归档:包含UniProt参考蛋白质组(Proteomes)2021年4月版(2021_04)及其OMArk质量评估结果。该归档同时包含运行OMArk所用的OMAmer分析结果(OMAmerResults)、OMArk质量评估结果(OMArk_Results)以及BUSCO完整性评估结果(BUSCO)。



