DFT+NEGF charge-transport dataset for double-stranded DNA duplexes (4-16 bp)
收藏资源简介:
Ab-initio charge-transport data for double-stranded DNA, computed with density functional theory (DFT) and the non-equilibrium Green's function (NEGF) formalism, for training and evaluating machine-learning models of transport. Part of ICLR 2027 conference submission number 50206. Contents520 randomly generated duplexes of 4-8 base pairs (2,077 records) form the training set, and 8 duplexes of 12 and 16 base pairs (32 records) form a held-out, out-of-distribution set: ACTTGACAGTCA, ATATATATATAT, GGGGGGGGGGGG, GTCAGGATCTGA (12 bp) and GGATAGGCTTAGAATT, GGGGGGGGGGGGGGGG, GTCAGGATCTGACAGT, TCAGTGCTAAGTCATG (16 bp). Each duplex is computed at four electrode configurations: two coupling strengths (0.1 or 0.6 eV) and two attachment geometries. The left contact is always applied to the first primary base; the right contact is applied to the last primary base ("same") or its complement ("cross"). Two training duplexes (CGTAT and GCCTGG) lack three corrupted configurations. Method Idealized B-form structures were built with the nucleic acid builder (nab). Gaussian 16 single-point calculations used B3LYP/6-31G(d,p) with a polarizable continuum model for water and a net charge of -2(N_bp - 1). The Fock (F) and overlap (S) matrices were orthogonalized as H = S^-1/2 F S^-1/2, and transmission T(E), density of states DOS(E) and per-atom local density of states were computed with a wide-band contact self-energy applied to every atomic orbital of the contacted residue. The energy grid spans the HOMO +/- 1 eV in 0.01 eV steps (201 points), referenced per sequence. Files:- transport.h5 (1.4 GB): every record's energy grid, T, DOS, per-atom LDOS, atom table (elements, names, residues, coordinates), Gaussian input text, contact atoms and residues, coupling, and a train/held-out split label. Units, conventions and the contact model are stored as file attributes.- matrices.h5 (13.8 GB, split equally into 10 parts): the Fock (Hartree) and overlap matrices for all 528 sequences with transport records, with a per-row orbital map (atom index, shell, component S/PX/PY/PZ/DXX..DYZ, and Gaussian's own basis-function type code). The map for the 8 held-out sequences was verified against Gaussian's matrix-element output. Because of its size it is uploaded as ten parts, matrices.h5.part0 ... part9; reassemble with: cat matrices.h5.part? > matrices.h5 (then check with sha256sum -c SHA256SUMS).- geom_cache.tar(500 KB): per-sequence X3DNA-DSSR rigid-body geometry caches (training and held-out) used by the models' optional geometry channel.- SHA256SUMS(238 B): checksums for all files. UsageThe accompanying code repository (https://anonymous.4open.science/r/G3NAT/) reads these files: DNADataset/import_hdf5.py converts transport.h5 into the per-record format used for training, and DNADataset/README.md in that repository documents every field. A companion record holds the trained model checkpoints (https://doi.org/10.5281/zenodo.22964131). LimitationsStructures are idealized B-DNA with no molecular dynamics or per-sequence relaxation; geometry varies only through base identity, so conformational and electronic effects are not separable in this dataset. Transport is coherent and zero-bias only.



