遇见数据集

CyanoCIN v1.1 : Curated Database of Cyanobacterial RiPP Precursor Peptides

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

CyanoCIN v1.1 is a curated database of cyanobacterial ribosomally synthesized and post-translationally modified peptide (RiPP) precursor peptides. This deposit contains the complete curated dataset underlying the v1.1 release, the relational database schema, SQL patches and cleanup scripts, benchmarking code, and supplementary benchmark data. Database: http://bgagenomics.iicb.res.in/ Contents File Description CyanoCIN_v1.1_data_table.xlsx Complete curated dataset containing 10,258 entries across 113 columns. This is the source table from which the database is built. CyanoCIN_v1.1_data_table.csv The complete curated dataset in CSV format for long-term accessibility and reuse. Table_S1_pipeline_benchmark.xlsx Supplementary Table S1 containing benchmark definitions, per-class breakdowns, and per-record outcomes. validation.py Benchmarking script used to reproduce the recovery and validation results reported in the accompanying paper. requirements.txt Python package requirements for running validation.py. Fig_S2_source_data.csv Strain counts per genus used as source data for Supplementary Figure 2. schema_and_patches.rar Relational database schema and SQL scripts associated with the CyanoCIN v1.1 release. schema_and_patches.rar File Description relational_schema_SupplementaryFigure1.png Relational schema diagram showing the PEPTIDE, GENOME, CONTIG, BGC_CLUSTER, PREDICTION_TOOL, and PEPTIDE_PREDICTION tables and their relationships. ripp_schema_and_loader.py Defines the relational database tables and loads the curated data table into them. Set RIPP_DB_URL before execution. cyanocin_v1.1_patch.sql Adds the four v1.1 fields to the normalized peptide table created by the loader. cyanocin_v1.1_patch_Peptide_v3.sql Adds the same four fields to Peptide_v3, the flat table used by the CyanoCIN web database. cyanocin_v1.1_cleanup_Peptide_v3.sql Removes placeholder values and redundant columns from Peptide_v3, corresponding to the cleanup applied to the v1.1 dataset. column_cleanup_report.csv Documents the columns removed or merged between CyanoCIN v1.0 and v1.1 and the rationale for each change. Fields Added in v1.1 Four fields were added to the CyanoCIN entries in v1.1. No v1.0 values were altered, except for the removal of the placeholder GenBank identifier (0) from 31 entries. Field Description RiPP_secondary_class Additional RiPP classes supported by the same precursor, separated by ; . Filled for 1,019 entries. Supporting_evidence Mining tools and curated sources supporting the entry, separated by ; . Tool_support_count Number of the three mining tools (antiSMASH v6.0.1, BAGEL4, and RiPPMiner-Genome) supporting the precursor, ranging from 0 to 3. Literature_DOI DOI of the source publication. Filled for 2,965 entries from 50 publications. Reproducing the Benchmark Install the required Python packages and run: pip install-r requirements.txt python validation.py CyanoCIN_v1.1_data_table.xlsx results/ The benchmark requires a few minutes to complete, primarily because of the pairwise sequence alignments used to score partial matches. The script generates the following files in the specified output directory: validation_summary.csv — headline benchmark results validation_per_class.csv — recovery results by RiPP main class validation_corroboration.csv — per-tool corroboration counts validation_per_record.csv — outcome for each literature record validation_results.pkl — Python pickle containing the corresponding result objects The benchmark reproduces the following figures reported in the accompanying paper: 2,965 literature-curated records, of which 2,507 are recoverable in principle 1,313 recovered as exact sequence matches (52.4%) and 167 as partial matches (6.7%) 1,480 records recovered overall (59.0%) Per-tool recall: BAGEL4 44.4%, antiSMASH v6.0.1 29.5%, and RiPPMiner-Genome 10.1% 92 of the 93 recoverable records with a reported mature product were recovered exactly 5,116 unique precursors predicted by the pipeline, of which 3,830 did not match any previously reported record Citation Please cite both the accompanying paper in Microbiology Spectrum (currently in revision) and this Zenodo deposit when using the CyanoCIN v1.1 dataset or database. Each CyanoCIN release is archived separately under its corresponding version number. Licence This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) licence, unless otherwise specified by an accompanying licence file.

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务