遇见数据集

Results and Analysis Dataset accompanying "OrthoSeq: A Design Workflow for Thermodynamically Orthogonal DNA Sequence-Pair Libraries"

收藏
Zenodo2026-07-29 更新2026-08-01 收录
官方服务:

资源简介:

OrthoSeq Benchmark Data This archive contains the benchmark data associated with the manuscript “OrthoSeq: A Design Workflow for Thermodynamically Orthogonal DNA Sequence-Pair Libraries.” The most directly reusable files are in extracted_libraries/, which contains ready-to-use Excel workbooks with the largest orthogonal sequence-pair libraries found for each reported benchmark condition. The archive also includes the full benchmark outputs used to derive those libraries. Archive scope The benchmark data are organized into two regimes: short_seq/ contains benchmarks for core binding-domain lengths 4, 5, 6, and 7 nt. Finite frozen candidate datasets were generated in advance, and all algorithms were evaluated on the same saved dataset bundles. long_seq/ contains benchmarks for core binding-domain lengths 8, 9, 10, 12, 14, 16, 18, 20, and 25 nt. Benchmark conditions were generated from preparation TOML files, and naive and hybrid searches were run live under a fixed budget of 10 million NUPACK calls. Both regimes include conditions with and without a 5′ TTTT extension. The main-manuscript long-sequence benchmark uses the 37 °C datasets. A 25 °C repeat is reported in the Supporting Information. Archive contents extracted_libraries/: ready-to-use Excel workbooks containing the largest selected sequence libraries for each reported condition. full_benchmark_results/short_seq/: frozen short-sequence datasets, cached thermodynamic matrices, benchmark summaries, and complete result workbooks. full_benchmark_results/long_seq/: generated benchmark condition files and complete live-search outputs. seqwalk_comparison/: the three workbooks used for the Figure 5 comparison between thermodynamic search and seqwalk-derived libraries. README.md: the detailed archive layout, dataset-array definitions, indexing conventions, workbook schema, metadata-key reference, naming conventions, and relevant source-code locations. Selection of extracted libraries For each reported condition, the extracted workbook was selected from the benchmark run that produced the largest final orthogonal sequence-pair library, measured by the number of rows in found_pairs. These are the files to use first when selecting sequences to inspect or copy into another workflow. Each reduced workbook contains only the run_metadata and found_pairs sheets. Short-sequence libraries are grouped by temperature, flank condition, conflict probability, and core binding-domain length. Long-sequence libraries are grouped by temperature, flank condition, and core binding-domain length. Full benchmark workbooks Full benchmark workbooks typically contain the following sheets: run_metadata found_pairs selected_hh selected_hah selected_ahah search_progress validation Hybrid long-sequence workbooks additionally contain seed_pass_pairs, seed_hh, seed_hah, and seed_ahah. Short-sequence datasets Each short-sequence dataset directory contains a human-readable dataset.toml, a compressed dataset.npz, and benchmark workbooks under results/. The NumPy archive stores the full candidate pool, canonical pair identifiers, sequences and intended partners, on-target and self-folding energies, the on-target-filter mask, and cached off-target energy matrices for the filtered matrix subset. Matrix row or column i corresponds to matrix_global_pair_ids[i]. Two sequence pairs are treated as incompatible when any relevant off-target interaction falls below the selected off-target free-energy cutoff. The accompanying dataset.toml records the indexing conventions, dataset inputs, NUPACK conditions, and derived statistics in plain text. Long-sequence datasets The configs/generated/ folders contain the parameter files used to define the long-sequence runs, including a batch summary, one condition TOML and job wrapper per run, and submit_all.sh. These files are inputs rather than results. Each run output contains one full XLSX workbook, one on-target/off-target PDF, and one self-folding PDF. The XLSX workbook is the primary machine-readable result artifact. Some 37 °C long-sequence runs request search.initial_fresh_pair_count = 2500. In practice, the fixed 10-million-call budget was exhausted during the initial graph-aware search after approximately 2235 sequence pairs. Such a run appears in the extracted libraries only if it produced the largest final found_pairs count for its condition. seqwalk comparison The seqwalk_comparison/ folder contains: figure5_seqwalk_max_orthogonality_len16_n72_seed42.xlsx: the seqwalk-only arm used for Figure 5A. figure5_search_only_hybrid_len16_noflank_init450.xlsx: the benchmark-derived hybrid-search arm used for Figure 5B. figure5_hybrid_len16_noflank_seqwalk_k6_seed42.xlsx: the seqwalk-derived candidate pool followed by thermodynamic hybrid-search filtering, used for Figure 5C. The Figure 5B workbook is duplicated in this folder so that all three comparison arms can be inspected together. Terminology vertex_cover: graph-aware search naive or naive_search: naive search hybrid_offline: hybrid search on frozen short-sequence datasets hybrid_search: live hybrid search in the long-sequence benchmark offtarget_limit: off-target free-energy cutoff initial_fresh_pair_count: initial graph-aware search subset size vc_max_iterations: number of graph-aware search iterations Workbook metadata keys use the prefixes input.*, search.*, nupack.*, dataset.*, and artifact.*. The artifact.* values may contain absolute paths from the machine on which the benchmark was run. These paths are retained only as provenance records and are not required to interpret the archive. Relation to the manuscript The short-sequence data support the benchmark regime in which full conflict-graph construction was feasible. The long-sequence data support the fixed-budget live-search regime. The 37 °C long-sequence data correspond to the main benchmark. The 25 °C data correspond to the lower-temperature repeat in the Supporting Information. The seqwalk_comparison/ workbooks support the Figure 5 comparison. Funding and compute resources This work was supported in part by: the U.S. Department of Energy, Office of Science, Basic Energy Sciences, Biomolecular Materials Program, under Award No. DE-SC0024136; the German Research Foundation (Deutsche Forschungsgemeinschaft, DFG) Walter Benjamin Programme, project 553862611; the Dana-Farber Cancer Institute Claudia Adams Barr Program for Cancer Research; and the Korea–US Collaborative Research Fund (KUCRF), grant RS-2024-00468463. The O2 High Performance Compute Cluster, supported by the Research Computing Group at Harvard Medical School, was used to accelerate development of the evolutionary algorithm and the final large-scale parameter sweeps.

提供机构:
Zenodo
创建时间:
2026-07-29
二维码
社区交流群
二维码
科研交流群
商业服务