Fairness in histopathology AI — systematic review: extraction, audit, and verification artifacts (79-study corpus, manuscript v0.3.5)
收藏资源简介:
Machine-readable evidence and analysis code for a PRISMA-2020 systematic review of fairness in histopathology AI. Contents. (i) The dual-coded, adjudicated extraction table for the 62-study analytical corpus, with the raw multi-coder outputs and inter-coder agreement statistics; (ii) the screening, de-duplication and post-hoc eligibility-audit logs; (iii) per-coder and adjudicated QUADAS-2 risk-of-bias judgements for all 62 studies; (iv) the full consensus-endorsement coding pipeline — screened review pools, the pre-specified random sample (n=30, seed 20260913), the coding specification, two independent coder passes, disagreements and their adjudication, and tallies with Wilson 95% confidence intervals; (v) the 39-dataset demographic-metadata audit; and (vi) a claim-by-claim statistics register mapping every reported number to its source rows, together with the code that produces each value. Full texts of the reviewed papers are not included, for copyright reasons. Code is licensed MIT; data tables are licensed CC BY 4.0. A SHA-256 manifest (MANIFEST.sha256) covers every file in the package. This package is a working-tree snapshot of the manuscript as of 30 September 2026 (46 pp main + 28 pp supplementary). Version of 30 September 2026. Adds the 79-study update: the dual-coded extraction and risk-of-bias tables including the 17 added studies, the post-hoc coding directory, a disposition for each of the 114 advanced screening records, the resolution of five full-text candidates, and a sensitivity analysis with one internally inconsistent study removed. Full texts of the reviewed papers are not included. Manuscript v0.3.0 (30 September 2026). The 79-study tables are unchanged. This version records that the manuscript now uses those tables in the discussion and the PRISMA checklist: author geography 81.0% (64/79), East Asia 12 studies, pediatric cancer the primary focus of two studies; any fairness metric 27.8% (22/79), expected calibration error in three studies; overall risk of bias 9% low (7/79) and 72% unclear (57/79). Full texts of the reviewed papers are not included. Manuscript v0.3.1 (30 September 2026). Third-pass adjudication of all 79 studies. A correction was kept only when confidence was high and the quotation was an exact substring of the full text. No human re-check. Current counts: at least one bias category 67/79 (84.8%); any fairness metric 25/79 (31.6%); external validation 41/79 (51.9%); TCGA 36/79 (45.6%); overall risk of bias 6 low, 54 unclear, and 19 high. Mitigation grade labels are unchanged; feature-level harmonization remains the only Strong grade. These counts supersede the v0.3.0 figures in the previous paragraph. The manuscript is 46 pages and the supplementary materials are 28 pages. Full texts of the reviewed papers are not included. Manuscript v0.3.2 (30 September 2026). Leftover-passage alignment after the v0.3.1 coding correction. No registered statistic changed. Appendix S2 category counts, the Results mitigation-table key findings, the Nagpal class/label parenthetical, the claim that a current original-62 split reproduces the pre-correction figures, and the S1 venue footnote and per-study venue column were brought into line with rows 43–55. Headlines remain 67/79 any bias category, 25/79 (31.6%) any fairness metric, 41/79 (51.9%) external validation, and overall risk of bias 6 low, 54 unclear, 19 high. The manuscript is 46 pages and the supplementary materials are 28 pages. Full texts of the reviewed papers are not included. Manuscript v0.3.3 (30 September 2026). Writing alignment after the v0.3.2 leftover-passage audit. No registered statistic changed. Methods and Discussion name both category-audit steps; Discussion fairness-aware evidence is 3 positive and 2 mixed; Results treats the ASCO abstract polin2025mutation as the meeting abstract of lin2025contrastive; the algorithmic n=5 paragraph cites the five coded members; mixed collaborations are classified by first-listed region (64/79 North America or Europe unchanged); the abstract is about 274 words; Discussion limitations are numbered First–Fifth. Headlines remain 67/79 any bias category, 25/79 (31.6%) any fairness metric, 41/79 (51.9%) external validation, and overall risk of bias 6 low, 54 unclear, 19 high. The manuscript is 46 pages and the supplementary materials are 28 pages. Full texts of the reviewed papers are not included. Manuscript v0.3.4 (30 September 2026). Jev classification of selected manuscript passages, then a rewrite of the sentences those questions flagged. No registered statistic changed. Introduction no longer opens with "woven throughout" or treats fairness as a deployment prerequisite this review establishes. Discussion drops the contrastive not-X-it-is-Y sentence and recasts policy "should require" as "We propose". Conclusion drops the four-constituency bind-to-authorization paragraph and labels the checklist as a draft. Headlines remain 67/79 any bias category, 25/79 (31.6%) any fairness metric, 41/79 (51.9%) external validation, and overall risk of bias 6 low, 54 unclear, 19 high. The manuscript is 46 pages and the supplementary materials are 28 pages. Full texts of the reviewed papers are not included. Manuscript v0.3.5 (30 September 2026). Numerical-claim verification pass: 69 paper-specific claims checked against corpus full texts, flagged items adjudicated by hand. Corrections include the roschewitz Youden before/after (0.295 to 0.651), AIDA 78.82% to 75.82%, the vaidya near-ceiling example, komen2024batch removed from the AUROC citation, and scope fixes for suzuki, lin2025stain, FAIR-Path, lafarge, tampu and musa2026cross. Headlines remain 67/79 any bias category, 25/79 (31.6%) any fairness metric, 41/79 (51.9%) external validation, and overall risk of bias 6 low, 54 unclear, 19 high. The manuscript is 46 pages and the supplementary materials are 28 pages. Full texts of the reviewed papers are not included.



