遇见数据集

CheckAMG manuscript source data

收藏
Zenodo2026-09-26 更新2026-10-01 收录
官方服务:

资源简介:

Source and supplemental data files for the CheckAMG manuscript. This record holds the training data, test data, per gene annotation tables, and model checkpoint referenced as Supplemental Data 1 through 7 in the manuscript. The CheckAMG software databases used to run the tool itself are hosted in a separate record, the CheckAMG database (10.5281/zenodo.21776005). Contents Seven gzipped tar archives (SDxx.tar.gz), numbered as they are cited in the manuscript (Supplemental Data 1 through 7). Each archive contains its own README.txt. Supplemental Data 1 (SD01.tar.gz, 3.3 GB) Amino acid sequences (FASTA) of the training and test datasets used to develop the CheckAMG viral origin classifier (LightGBM). train.faa: 13,999,957 proteins test_input.faa: 960,240 proteins, source composition matched to the training set test_equal_pos.faa: 383,948 proteins, 50% virus, 25% host, 25% MGE test_equal_source.faa: 191,832 proteins, equal virus, host, and MGE test_half_virus_host.faa: 635,737 proteins, 50% virus and 50% host, no MGE test_virus_enriched.faa: 427,977 proteins, 75% virus test_host_enriched.faa: 313,192 proteins, 75% host test_mge_enriched.faa: 203,839 proteins, 75% MGE test_near_all_virus.faa: 846,398 proteins, 90% virus test_near_all_host.faa: 667,539 proteins, 90% host test_provirus.faa: 75,260 proteins from chromosomal contigs carrying both host and integrated provirus regions Percentages are the target compositions of each test set by the source (virus, host, or MGE) of the contig each protein was predicted from. Supplemental Data 2 (SD02.tar.gz, 832 MB) Per protein feature tables (Apache Parquet) for the same training and test datasets as Supplemental Data 1. Sequences for these proteins are in Supplemental Data 1. One row per protein and 64 columns: protein, contig, and genome identifiers, gene coordinates and contig context, Pfam, KEGG, and PHROG V-scores and VL-scores, contig and window averages of those scores, distances to the nearest viral and MGE genes, the LightGBM viral probability and the resulting viral origin confidence from the model used by CheckAMG annotate at the time of labeling (outdated, not the final model results), and the source and classification labels used for training and evaluation. train.parquet: 13,999,957 rows test_input.parquet: 960,240 rows, source composition matched to the training set test_equal_pos.parquet: 383,948 rows, 50% virus, 25% host, 25% MGE test_equal_source.parquet: 191,832 rows, equal virus, host, and MGE test_half_virus_host.parquet: 635,737 rows, 50% virus and 50% host, no MGE test_virus_enriched.parquet: 427,977 rows, 75% virus test_host_enriched.parquet: 313,192 rows, 75% host test_mge_enriched.parquet: 203,839 rows, 75% MGE test_near_all_virus.parquet: 846,398 rows, 90% virus test_near_all_host.parquet: 667,539 rows, 90% host test_provirus.parquet: 75,260 rows, proteins from chromosomal contigs carrying both host and integrated provirus regions Percentages are the target compositions of each test set by the source (virus, host, or MGE) of the contig each protein was predicted from. Supplemental Data 3 (SD03.tar.gz, 148 MB) Raw, unfiltered per gene output (combined from all per-sample runs into single Apache Parquet tables) from each tool benchmarked for AMG prediction. Columns shared by all three files identify the gene, its coordinates and frame, its scaffold, and the sample, source, and ecosystem it came from. The remaining columns are the native output fields of each tool. checkamg_results_raw.parquet: CheckAMG, 4,316,180 rows by 32 columns dramv_results_raw.parquet: DRAM-V, 10,170,394 rows by 24 columns vibrant_results_raw.parquet: VIBRANT, 16,362 rows by 17 columns Supplemental Data 4 (SD04.tar.gz, 10.8 GB) amg_all_annotations.parquet: all database hits recovered from the annotation searches of CheckAMG, DRAM-V, and VIBRANT, restricted to the benchmark genes that at least one of the three tools predicted as an AMG (903,000,986 rows by 11 columns, about 10.8 GB uncompressed). Hits are limited to those with an e-value of 1e-3 or lower and a bitscore of 30 or higher. One row per hit. Columns: gene, source, sample, ecosystem, tool, database, hit_id, hit_desc, bitscore, evalue, and final_annot. final_annot is true when the tool reported that hit in its final output for the gene (Supplemental Data 3), and false for hits that passed the search but were not retained. Supplemental Data 5 (SD05.tar.gz, 56.8 GB) Training and test data for CheckAMG-PST: protein sequences, per protein labels, and the ESM2 (esm2_t30_150M) embeddings in PST graph format that the model takes as input. Test set names match those in Supplemental Data 1 and 2. Protein counts are slightly lower than in Supplemental Data 1 and 2 because contigs with only one predicted protein were removed. train_data train_ptns.faa: 13,942,367 proteins train_ptns.parquet: labels for those proteins (Contig, Protein, frame, gene_number, viral, AVG) checkAMG_train_esm2_t30_150M.graphfmt.h5: ESM2 embeddings for train_ptns.faa test_data test_SETNAME_ptns.faa and test_SETNAME_ptns.parquet: the 10 held out test sets, in the same format as the training files (SETNAME is one of the 10 test set names used in Supplemental Data 1 and 2, for example test_input, test_equal_pos, test_provirus) checkAMG_test_SETNAME_esm2_t30_150M.graphfmt.h5: ESM2 embeddings for each test set checkAMG_test_combined.ptns.faa and checkAMG_test_combined.h5: the 10 test sets concatenated (4,679,640 proteins) and their combined embeddings Supplemental Data 6 (SD06.tar.gz, 180 MB) checkAMG-PST_TL-P__large_5.20260714.ckpt: the trained CheckAMG-PST model checkpoint distributed with the CheckAMG de novo database v1.1 (build 20260714). It was fine tuned from the pre-trained PST-TL-P__large checkpoint on the data in Supplemental Data 5, and CheckAMG de novo uses it to generate the protein embeddings that its AVG predictions are based on. Supplemental Data 7 (SD07.tar.gz, 90 MB) avg_proteins_soil_gut.faa: amino acid sequences of all 541,199 proteins predicted as AVGs by CheckAMG annotate, CheckAMG de novo, or both, across the MetaVR soil and gut datasets. Of these, 289,167 are from soil and 252,032 from gut, 351,814 were predicted by annotate and 484,165 by de novo, and 294,780 were predicted by both. Sequences and headers are unmodified pyrodigal-GV output from those CheckAMG runs, so each header holds the protein identifier followed by the gene coordinates, strand, and gene calling attributes. Per protein predictions and annotations for these proteins are reported in the accompanying supplemental tables. Citation If you use this data, please cite the CheckAMG manuscript (citation to be added on publication) and, if applicable, this Zenodo record.

提供机构:
Zenodo
创建时间:
2026-09-26
二维码
社区交流群
二维码
科研交流群
商业服务