SCN5A R104Q: guide design, off-target scan, and ClinVar functional-evidence census
收藏资源简介:
Data tables behind two preprints on adenine base editor guide design for SCN5A p.Arg104Gln and on the retrievability of deposited functional evidence in ClinVar. Computational predictions and database measurements only; no experimental validation. Version 2, 7 August 2026. This version corrects one wrong data file, documents a second that is known to be defective and is NOT corrected here, and — the serious part — deposits the tables that seven of the ten papers stated were deposited here and were not. What changed, measured against the published version 1 rather than asserted. Version 1 held 52 files; version 2 holds 107. Fifty-five files are new and none was removed. Of the original 52, fifty are byte-identical to version 1; the two that changed are MS_TABLE3_LEAD_PROTEIN_CHANGING.csv, which was wrong and is corrected, and README.md, which gained the contents listing and correction notes for everything added. The new files cover paper 10 (7 files), paper 5 (5), paper 6 (8), paper 7 (5), paper 2 (10), paper 11 (8), paper 8 (8) and paper 3 (1), plus three route-10 derived tables and the two protocol documents paper 11's section 3 turns on. The pattern behind most of it, stated plainly because it is the least flattering thing here. All ten deposited papers were audited and eight of them named tables as available in this archive that this archive did not hold — papers 1, 3, 5, 6, 7, 8, 9 and 10. Only papers 2 and 4 were clean. Every one of those statements was checkable and none had been checked. They were found by auditing all eleven data availability statements one at a time between 6 and 7 August 2026. Where the underlying analysis was regenerable it was regenerated from public inputs and verified quantity by quantity against what each paper prints; where it was not, the papers now say so rather than claiming otherwise. Three papers could not be repaired by regeneration and are corrected by declaration instead. Papers 1, 3 and 9 name material that does not exist anywhere — a Ssym/S669 benchmark table, three parser and reconciliation scripts, and a 151-variant scoring run with its calibration gate and tier list. All were searched for across the project tree and the compute host on 7 August 2026 and no copy of any of them was found. Version 2 of each paper names what is deposited and states plainly what is absent, which is the only honest repair available. One file of the four paper 3 named was recoverable, CENSUS_DISCREPANCY_RESOLVED.md, and it is deposited here. The paper-10 defect, added to this note 6 August 2026 at 22:32, and it is the one to read first. Paper 10 (10.5281/zenodo.21799871) stated that its derived tables were deposited here and named four of them: the pairwise geometry table, the per-residue neighbour count table, the per-residue sigmoid weight table, and the recompute output. Version 1 held none of the four, and held no paper-10 file of any kind — of the ten papers version 1 supported, paper 10 was the only one with nothing here. That is a false statement in a published, citable record, and it is a worse class of defect than a wrong row count. It could not be fixed by locating the missing files: the code that made them is not in the project tree and not on the compute host, both were searched on 6 August, and there is no R installation on either. Three of the four tables are pure geometry over the public 8VYJ coordinates and were therefore regenerated from a script written the same evening and deposited with it, then checked against every number paper 10 prints — thirty-five of thirty-seven reproduce exactly. The fourth, the recompute output, needs the kroncke-lab/Bayes_BrS1_Penetrance R pipeline and an expectation-maximisation refit, and is still absent; version 2 of paper 10 says so rather than claiming otherwise. Full account: P10_REGENERATION_NOTE.md inside the archive, and the correction block at the end of the archive README. One further finding travelled with the reproduction and is a defect in paper 10 rather than in this deposit: under the definition paper 10 states in its own table, the domain contains four salt bridges, not two. It is corrected in version 2 of that paper. No conclusion moves. What was wrong. MS_TABLE3_LEAD_PROTEIN_CHANGING.csv as deposited on 5 August 2026 held 16 data rows. The paper that cites it states the count is 22 and prints all 22 in its own Table 3. A reader who downloaded this archive to check that paper found the paper's supporting file contradicting it, with nothing beside it to say which was right. The paper was right and the data file was wrong, and it is worth saying plainly that the defect looked worse than it was: it reads as though an off-target count was inflated after the fact, and it was not. Why it happened. The 5 August file was built from the worst_consequence column of ABE_OFFTARGET_SCAN.csv. That column holds the original consequence calls, which were shifted by one residue in 133 rows and were superseded by the recomputed_worst column of ABE_CONSEQUENCE_RECHECK.csv. The paper says so explicitly and states that its Table 3 comes from the recheck file. The summary CSV was never rebuilt to follow the recount. On the lead guide's 26 coding sites the original column reads 16 missense, 9 synonymous and 1 uncalled; the corrected column reads 22 missense and 4 synonymous. Nothing was withheld in version 1. ABE_CONSEQUENCE_RECHECK.csv, the file the paper points at for full detail, was present and complete at 22,717 rows in version 1 and is unchanged in version 2 was present at 22,717 rows in version 1 and is unchanged in version 2. Both the wrong summary and the right detail were always in the same archive. Correction to the paragraph above, 6 August 2026 evening, and it is the reason this version note is not a clean bill of health. The word "complete" was wrong when it was written. ABE_CONSEQUENCE_RECHECK.csv truncates its recomputed_aa field at exactly 1,500 characters. Fifty-eight of its 668 call-carrying rows sit at the cap and end mid-token — the first stops at …NRXN3:XM_047431956.1: with no residue and no consequence class — and no row exceeds it. The file prints 7,137 of the 8,843 transcript-level consequence calls it should carry; 1,706, or 19.3 percent, are absent, across 28 genes. The row count of 22,717 is right and says nothing about the defect, because the truncation is inside one field of 58 rows and the truncated rows do not look truncated. This file is unchanged in version 2 and is therefore still defective in version 2. It is not repaired here because repairing it requires the script that produced it, and that script is not in the project tree; there is no second copy of the CSV, so the text of the missing calls is gone for good. What is established is which calls are missing: the generating rule was validated against the 610 rows the cap never touched, where it predicts 5,032 calls and the file carries exactly those 5,032, with none missing and none extra. The clean repair is to regenerate this table without the cap and deposit it as a version 3. Until then, any count taken from this file is a count over the 7,137 calls it prints. Version 2 of paper 2 (10.5281/zenodo.21799850) states this in its section 3.4.8, in limitation 6.1.11 and in its data availability section, so a reader arriving from that paper is warned; a reader arriving at this deposit directly is warned by this note and by the archive README. Why version 2 is still worth publishing with a known-defective file in it. The defect it fixes — a 16-row summary table contradicting the paper that cites it — is the one a reader hits first and cannot diagnose. The defect it does not fix is now documented in three places instead of zero. Publishing the correctable half and naming the uncorrectable half is a better record than waiting for a script that may never be rewritten. How the replacement was derived, so a reader can check it rather than trust it. Lead guide is the 655 sites with spg_route_editable == True; coding is the 26 of those with region == 'exonic_CDS'; protein-changing is the 22 of those whose recomputed_worst is missense. All 22 rows were then compared cell by cell against the paper's printed Table 3 — gene, consequence, mismatch count, seed mismatch count, heart left-ventricle TPM — and all 22 agree, in the same order. The script that does it and the selection rule it applies are stated in the archive's README. Two smaller changes travel with it and are recorded rather than slipped in. The representative transcript for MSH6 moves from NM_001406809.1 to NM_000179.3, because the paper's stated rule is the lowest-numbered RefSeq transcript carrying the consequence and NM_000179.3 is lower; the amino-acid change is p.V263A either way. And amino-acid changes are now written one-letter rather than three-letter, because that is what regenerating from ABE_CONSEQUENCE_RECHECK.csv produces, and the three-letter form was a fingerprint of the superseded scan table. What this does not change. No count, no conclusion and no claim in any of the ten papers moves. The README of the archive carries the same note with the full derivation. Version 1 is not withdrawn. It remains resolvable at 10.5281/zenodo.21799234 and readers may hold its 16-row file. It is superseded, not deleted, and this note exists so that anyone comparing the two can see exactly why. Everything in this deposit is computation over public reference data. No cell was edited and no current was recorded. Every therapeutic reading of these tables is a hypothesis, not a demonstrated result. Added 7 August 2026 — five files for paper 5, for the same reason as paper 10's seven Paper 5 (10.5281/zenodo.21799861) says its recomputed statistics, paralogue alignment counts and baseline-recalibration figures are deposited here. Version 1 held the paralogue alignment counts and neither of the other two. Same defect class as paper 10, one degree less severe, and it is a false statement in a published record rather than a wrong number in a table. Unlike paper 10's geometry, both missing tables were regenerable and both were regenerated. They are arithmetic over published summary values from O'Neill 2022 Supplementary Table 1 (PMID 35305865, also bioRxiv doi:10.1101/2021.09.22.461398) and Gütter 2013 Table 3 (PMID 23805106). Paper 5's own Methods record that the recomputation "was done by hand" with no software version, so no script was lost. p5_regen_statistics.py, written 7 August 2026 from the stated method with the Python standard library only, reproduces all 32 quantities paper 5 prints. Added: P5_REGENERATION_NOTE.md — provenance, method, what is deliberately absent. Start here P5_RECOMPUTED_STATISTICS.csv — every recomputed comparison, with the Welch alternative alongside the z form P5_BASELINE_RECALIBRATION.csv — the dominant-negative baseline recalibration over all 51 paired variants of Table S1, summary block plus per-variant block p5_regen_statistics.py — the script producing both P5_VERIFICATION_OUTPUT.txt — its verification print-out, every reproduced quantity beside paper 5's printed value Two things are deliberately not here and are stated in paper 5 rather than left to be discovered. The para-SAME sweep across 5,559 ClinVar variants in the nine Naᵥ paralogues is not deposited and was never among the tables the statement named. And P5_BASELINE_RECALIBRATION.csv carries raw per-cell values only for the six variants paper 5 already prints, marking the other 45 rows raw_values_withheld while keeping their derived flags, because O'Neill's Supplementary Table 1 is a third party's table this archive should not redistribute in full. Every quantity paper 5 prints stays checkable from the deposited flags and summary block. One thing the regeneration found, recorded rather than buried. Paper 5's Gütter differences use the convention variant minus wild type; computing them the other way gives every magnitude right and every sign inverted. The deposited table follows the paper's convention and says so. No number in paper 5 changes. Added 7 August 2026 — thirteen files for papers 6 and 7, for the same reason as papers 10 and 5 Paper 6 (10.5281/zenodo.21799863) says its per-sample junction and transcript classifications, its transcript biotype table and its headroom calculations are deposited here. Paper 7 (10.5281/zenodo.21799865) says its per-study power-calculation numbers and its dimer-arithmetic solutions for x are deposited here. Version 1 held not one file belonging to either paper. Same defect class as papers 10 and 5, and as total as paper 10's: false statements in published records. Paper 6's fix is a recovery, and saying otherwise would overstate it. Three of its four tables were written by the original pipeline on 4 August 2026 and were never lost, only never deposited. They are added under their original names. The fourth, the headroom table, also existed — carrying the retired 34.1 baseline and 50 comparator that paper 6's own correction section withdraws — so it was recomputed from 31.3, 45.8 and 100 and deposited in corrected form. The retired file is neither deposited nor deleted. Every quantity paper 6 prints was then recomputed from the deposited files and 116 of 119 reproduce at the printed precision. Added for paper 6: P6_REGENERATION_NOTE.md — provenance, verification, and what is not checkable. Start here SCN5A_SPLICING_MEASUREMENT.csv — the per-sample junction and transcript classifications, 2,942 rows SCN5A_SPLICING_SUMMARY.csv — the four headline group statistics jx_classification.csv — every annotated junction of both genes, classified ensembl_tx_biotypes.csv — the transcript biotype table, 21 SCN5A and 30 SCN1A transcripts UPREGULATION_HEADROOM_CORRECTED.csv — the headroom calculations on the corrected anchors p6_verify_and_headroom.py — the script that writes the headroom table and verifies the rest P6_VERIFICATION_OUTPUT.txt — its verification print-out Paper 7's fix is a true regeneration. Nothing of its calculation survived anywhere, and nothing needed to: every input is a published summary statistic and the method is fully stated in the paper. Both tables were rebuilt from first principles and 37 of the 39 quantities checked reproduce at the printed precision, including every solved x, both cross-study z-tests and all three 8VYJ distances recomputed from the public coordinates. Added for paper 7: P7_REGENERATION_NOTE.md — provenance, verification, and what could not be regenerated. Start here P7_POWER_INPUTS.csv — the per-study numbers, 17 rows with source, dispersion type, n and PMID P7_DIMER_ARITHMETIC.csv — the solutions for x, seven variants, with each mechanism's prediction p7_power_and_dimer.py — the script that writes both and recomputes the structural distances P7_VERIFICATION_OUTPUT.txt — its verification print-out One finding travelled with paper 6's reproduction and is a defect in the paper rather than in this deposit. Paper 6's Mann-Whitney p of 1.6e-160 is attributed to the junction comparison and is in fact the transcript comparison; the junction comparison gives 4.0e-219 with tie correction and 1.3e-168 without. The error understated paper 6's own result. Version 2 of that paper states both figures with the comparison each belongs to. Two smaller items — a rounding propagation in the 308-fold gap, which is 310.9-fold from the raw counts, and a last-digit slip in one 90th percentile — are recorded in the same place. Paper 7's two non-reproducing quantities are both rounding propagations and are recorded in version 2 of that paper. What is deliberately not here. Four claims in paper 6 rest on the GTEx v10 junction file, which was streamed rather than stored, so they are not checkable against any deposited file and version 2 of paper 6 says so. Paper 7's eighteen PubMed query strings were never recorded verbatim, so its retrieval counts cannot be re-derived, and version 2 of that paper says so. And O'Neill 2022's Supplementary Table 1 is not deposited in full, for the same reason P5_BASELINE_RECALIBRATION.csv withholds its per-variant raw values: it is a third party's supplementary material. The four rows paper 7 uses are quoted inside p7_power_and_dimer.py, and paper 7 prints all four in its own text. Version 2 addendum, 7 August 2026: paper 8, the last unaudited data availability statement, and it is the worst-notated case in the set. Lead with the negative. Paper 8 named four kinds of table as deposited here and one of them — the shell residue lists — was not. Worse, and in a place nobody was checking, the italic data line closing its Part 2 named six CSV files by name, and five of the six are not in this archive, are not in the project that produced them, and are not recoverable: POS104_SIDECHAIN_RSA.csv, POS104_MIN_ENERGETICS.csv, POSITION_SPECIFICITY_CONTROL.csv, TRP_EXPOSURE_VS_PHENOTYPE.csv and ctrl_energetics.csv. Naming a file makes a statement checkable; it does not make it true. A third gap nobody had claimed. Paper 8's Part 1 — the 1,008-model Rosetta and OpenMM analysis behind its sections 2 through 6a — has no file in this archive at all. HYDROPHOBIC_PATCH_ANALYSIS.csv was tested against Part 1's printed tables under every plausible reading of its 56 columns and reproduces none of them; it is a different run, and is the source of Part 2's Independent run column instead. What was added: eight files. The shell residue lists, the Part 5 site geometry, the Part 3 linear-motif scan, the Part 4 sequon scan, a per-residue Shrake–Rupley accessibility reference over all 1,395 modelled residues of 8VYJ chain A, the script that writes all five from public inputs, its verification output and its provenance note. 43 of 45 quantities reproduce exactly, including the shell lists themselves, the 50.7 Å Arg104-to-Tyr1767 separation, the zero shell overlap, all seven motif classes, the eight-RxxL sanity count, the four asparagines and zero sequons, and Arg104's 13.4 percent accessibility. The two that do not are in one sentence of Part 2's validation paragraph and are named in version 2 of paper 8 rather than adjusted. And three route-10 derived tables plus two protocol documents, for paper 11. Paper 11 named MOG1_GTEX_EXPRESSION.csv, MOG1_CELLTYPE_EXPRESSION.csv and MOG1_NAV15_PROTEIN_STOICHIOMETRY.csv in prose and none was here; every figure in its stoichiometry and regional-contrast paragraphs reproduces from the last of those alone. EXPERIMENT_PROTOCOLS.md and its cost companion ASSAY_COST_TIMELINE.csv were added because paper 11's section 3 turns on a single sentence of the first, and until now a reader could not check the sentence the paper's whole two-layer argument rests on. One file that could have been added and deliberately was not. SESSION_ARCHIVE_20260804/data/MODALITY_MAP.csv, which paper 11's section 5.5 cites, states the coupled-gating route's ceiling as "34 -> up to 50%" — both of them this project's retired anchors. Depositing a file that prints them without their correction beside them would propagate the exact error this project has already had to chase four times. Paper 11 restates that row on 31.3 and 45.8. Nothing in this archive was removed or altered by this rebuild except README.md, which gained two contents sections, two rows, a correction note and a corrected file count.



