遇见数据集

Results and Analysis Dataset accompanying "OrthoSeq: A Design Workflow for Thermodynamically Orthogonal DNA Sequence-Pair Libraries"

收藏
Zenodo2026-07-20 更新2026-08-01 收录
官方服务:

资源简介:

OrthoSeq Benchmark Data This archive contains the benchmark data associated with the manuscript “OrthoSeq: A Design Workflow for Thermodynamically Orthogonal DNA Sequence-Pair Libraries.” It includes the full benchmark outputs and the reduced sequence libraries used for the Supporting Information. Archive scope The benchmark data are organized into two regimes: short_seq/ contains benchmarks for core binding-domain lengths 4, 5, 6, and 7 nt. Finite frozen candidate datasets were generated in advance and all algorithms were evaluated on the same saved dataset bundles. long_seq/ contains benchmarks for core binding-domain lengths 8, 9, 10, 12, 14, 16, 18, 20, and 25 nt. Benchmark conditions were generated from preparation TOML files, and naive and hybrid searches were run live under a fixed budget of 10 million NUPACK calls. Both regimes include conditions with and without a 5′ TTTT extension. The main-manuscript long-sequence benchmark uses the 37 °C datasets; a 25 °C repeat is reported in the Supporting Information. Archive contents full_benchmark_results/short_seq/: frozen short-sequence datasets, cached thermodynamic matrices, benchmark summaries, and complete result workbooks. full_benchmark_results/long_seq/: generated benchmark condition files and complete live-search outputs. extracted_libraries/: reduced workbooks containing the selected sequence libraries for each reported condition. seqwalk_comparison/: the three workbooks used for the Figure 5 comparison between thermodynamic search and seqwalk-derived libraries. README.md: the detailed archive layout, dataset-array definitions, indexing conventions, workbook schema, metadata-key reference, naming conventions, and relevant source-code locations. Selection of extracted libraries For each reported condition, the extracted workbook was selected from the run that produced the largest final orthogonal sequence-pair library, measured by the number of rows in found_pairs. Each reduced workbook contains only the run_metadata and found_pairs sheets. Short-sequence libraries are grouped by temperature, flank condition, conflict probability, and core binding-domain length. Long-sequence libraries are grouped by temperature, flank condition, and core binding-domain length. Full benchmark workbooks Full benchmark workbooks typically contain the following sheets: run_metadata found_pairs selected_hh selected_hah selected_ahah search_progress validation Hybrid long-sequence workbooks additionally contain seed_pass_pairs, seed_hh, seed_hah, and seed_ahah. Short-sequence datasets Each short-sequence dataset directory contains a human-readable dataset.toml, a compressed dataset.npz, and benchmark workbooks under results/. The NumPy archive stores the full candidate pool, canonical pair identifiers, sequences and intended partners, on-target and self-structure energies, the on-target-filter mask, and cached off-target energy matrices for the filtered matrix subset. Matrix row or column i corresponds to matrix_global_pair_ids[i]. Two sequence pairs are treated as incompatible when any relevant off-target interaction falls below the selected off-target free-energy cutoff. The accompanying dataset.toml records the indexing conventions, dataset inputs, NUPACK conditions, and derived statistics in plain text. Long-sequence datasets The configs/generated/ folders contain the parameter files used to define the long-sequence runs, including a batch summary, one condition TOML and job wrapper per run, and submit_all.sh. These files are inputs rather than results. Each run output contains one full XLSX workbook, one on-target/off-target PDF, and one secondary-structure PDF. The XLSX workbook is the primary machine-readable result artifact. Some 37 °C long-sequence runs request search.initial_fresh_pair_count = 2500. In practice, the fixed 10-million-call budget was exhausted during the initial graph-aware search after approximately 2235 sequence pairs. Such a run appears in the extracted libraries only if it produced the largest final found_pairs count for its condition. Seqwalk comparison The seqwalk_comparison/ folder contains: figure5_seqwalk_max_orthogonality_len16_n72_seed42.xlsx: the seqwalk-only arm used for Figure 5A. figure5_search_only_hybrid_len16_noflank_init450.xlsx: the benchmark-derived hybrid-search arm used for Figure 5B. figure5_hybrid_len16_noflank_seqwalk_k6_seed42.xlsx: the seqwalk-derived candidate pool followed by thermodynamic hybrid-search filtering, used for Figure 5C. The Figure 5B workbook is duplicated in this folder so that all three comparison arms can be inspected together. Terminology vertex_cover: graph-aware search naive or naive_search: naive search hybrid_offline: hybrid search on frozen short-sequence datasets hybrid_search: live hybrid search in the long-sequence benchmark offtarget_limit: off-target free-energy cutoff initial_fresh_pair_count: initial graph-aware search subset size vc_max_iterations: number of graph-aware search iterations Workbook metadata keys use the prefixes input.*, search.*, nupack.*, dataset.*, and artifact.*. The artifact.* values may contain absolute paths from the machine on which the benchmark was run; these paths are retained only as provenance records and are not required to interpret the archive. Relation to the manuscript The short-sequence data support the benchmark regime in which full conflict-graph construction was feasible. The long-sequence data support the fixed-budget live-search regime. The 37 °C long-sequence data correspond to the main benchmark. The 25 °C data correspond to the lower-temperature repeat in the Supporting Information. The seqwalk_comparison/ workbooks support the Figure 5 comparison. Funding and compute resources This work was supported in part by: the German Research Foundation (Deutsche Forschungsgemeinschaft, DFG) Walter Benjamin Programme, project 553862611; the Dana-Farber Cancer Institute Claudia Adams Barr Program for Cancer Research; and the Korea–US Collaborative Research Fund (KUCRF), grant RS-2024-00468463. The O2 High Performance Compute Cluster, supported by the Research Computing Group at Harvard Medical School, was used to accelerate development of the evolutionary algorithm and the final large-scale parameter sweeps.

提供机构:
Zenodo
创建时间:
2026-07-20
二维码
社区交流群
二维码
科研交流群
商业服务