VAPOR benchmark data, pretrained reward models, and ESM-2 logits cache for protein sequence optimization
收藏资源简介:
This dataset provides the benchmark data and pretrained assets used byVAPOR, a Python toolkit for Pareto-leading protein sequence optimizationwithout invoking a protein language model during search. The deposit contains three archives: 1. vapor_benchmark_data.tar.gz: Benchmark DMS datasets used for protein sequence optimization experiments, including core benchmarks and prepared ProteinGym landscapes. 2. vapor_pretrained_models.tar.gz: Pretrained reward model ensemble checkpoints used by VAPOR. 3. vapor_esm2_logits_cache.tar.gz: Cached ESM-2 per-position token logits used for offline PLM distillation. These caches avoid runtime PLM calls during sequence optimization. These files are intended to be extracted into the VAPOR repository rootor release package directory before running the command-line interface.The source code is available separately as part of the VAPOR softwarepackage.



