遇见数据集

Processed benchmark artifacts for coding and non-coding genomic fragment classification across flatfish genomes

收藏
Zenodo2026-06-12 更新2026-06-17 收录
官方服务:

资源简介:

This dataset contains processed benchmark artifacts supporting the manuscript “A reproducible benchmark with sequence-similarity auditing for coding and non-coding genomic fragment classification across flatfish genomes”. The archive provides benchmark materials for coding versus non-coding genomic fragment classification across annotated flatfish genomes. It includes split manifests, task label mappings, regime definitions, evaluation summaries, length-binned performance tables, multi-seed summaries, MMseqs2 sequence-similarity audit summaries, leakage-sensitivity summaries, cross-species transfer summaries and diagnostic-task summaries. The benchmark defines five core regimes that vary sequence construction, split unit and exact de-duplication scope while preserving the same source genomes and annotation-derived label framework. The principal benchmark regime is Exp4_cross_dedup, which combines non-concatenated genomic fragments, gene-level partitioning and cross-species exact sequence-level de-duplication. The included MMseqs2 audit summaries provide post hoc nearest-neighbour similarity profiles between training and test partitions. Source genome assemblies and annotations are publicly available from NCBI RefSeq under the accessions listed in the manuscript. Source code for benchmark generation, model fine-tuning, evaluation, figure generation and sequence-similarity auditing is maintained separately in the associated code repository.

提供机构:
Zenodo
创建时间:
2026-06-12
二维码
社区交流群
二维码
科研交流群
商业服务