遇见数据集

Data for study "Selecting Denoisers for Frozen Pedestrian Detectors Under Gaussian Noise and JPEG Acquisition: Degradation Matching over Restoration Fidelity"

收藏
Zenodo2026-09-09 更新2026-10-01 收录
官方服务:

资源简介:

Basic information---------------------------1. Journal article: Selecting Denoisers for Frozen Pedestrian Detectors Under Gaussian Noise and JPEG Acquisition: Degradation Matching over Restoration Fidelity 2. DOI: 10.5281/zenodo.22643936 3. Contact information Name: Vo Thanh Kiet Institution: VSB – Technical University of Ostrava E-mail: kiet.vo.thanh.st@vsb.cz ORCID: https://orcid.org/0009-0002-3278-8755 4. Dataset publication date: 2026-09-07 5. Place of publication: Ostrava, Czechia ------------------------------------------------------------------6. Dataset Description================================================================================ This dataset contains the per-image fidelity and structure measurements, theper-cell detection results, the aggregated tables and pre-specified gatestatistics, the own-trained denoiser checkpoints together with the frozendetector used for the core grid, and the source code that produced theresults, tables and figures presented in the article. The study asks whichdenoiser should be placed in front of a pedestrian detector whose weights arefrozen -- a detector that is never fine-tuned under any restorer -- when theincoming frames carry additive Gaussian noise and are then stored by a lossyJoint Photographic Experts Group (JPEG) acquisition step. Its answer is thatdegradation matching beats restoration fidelity: a restorer trained on thedegradation actually present recovers more mean average precision atintersection-over-union 0.5 (mAP@50) than a blind, higher-fidelity restorerdoes, and the Structural Preservation Score (SPS) tracks that differencewhere the Peak Signal-to-Noise Ratio (PSNR) and the Structural SimilarityIndex Measure (SSIM) conceal it. The core grid is run on the People andPerson-Like Objects (PnPLO) pedestrian dataset; three external arms extend itto INRIA Person, CrowdHuman and CityPersons. Experimental design. The core grid uses the 235-image test split of PnPLO(944 / 160 / 235 train/val/test images) with a frozen YOLOv9-M detector thatscores mAP@50 = 0.745 on clean frames and falls to 0.695 / 0.542 / 0.399 atsigma = 10 / 20 / 30, the within-sigma baselines every denoiser must recoverfrom. The deployed stage order is fixed throughout: additive Gaussian noise,then a JPEG re-encode of the noisy frame at quality 95 (the acquisitionstep), then the denoiser, then the frozen detector. The controlled ablationlever is that JPEG stage, switched on and off with the restorer and detectorweights untouched, plus a quality sweep at q = 75 and q = 85 and a sigma = 50stress point outside the levels a blind restorer is trained for. Twelverestorers are evaluated: sigma-conditioned (SwinIR, SCUNet, FFDNet),level-blind (Restormer, PromptIR, NAFNet), classical or own-trained (BM3D,Gaussian filter, DnCNN, a convolutional autoencoder denoted CAE, and CAE-PSO,whose architecture hyperparameters are chosen by particle swarmoptimization), plus SCUNet-real, the real-noise SCUNet checkpoint appliedwithout level routing, used as the control that separates the "blind" labelfrom degradation match (arm3b in data/pnplo/arms_rows.csv). Public Gaussiancheckpoints ship at fixed levels, so SwinIR and SCUNet are applied at thenearest available level (15 / 25 / 25 at sigma = 10 / 20 / 30) while FFDNetand the classical methods take the true sigma. Three properties of everydenoised frame are recordedand tied to detection under one protocol: the within-sigma change in mAP@50(denoted d_mAP50), the structure scores SPS_grad and SPS_edge computed fromfrozen Sobel gradient and Canny edge maps, and the fidelity metrics PSNR andSSIM; latency is measured separately. The statistics are pre-specified: gatetests V1 to V4, within-sigma Spearman correlations with bootstrap intervals,an exact within-stratum permutation test, a paired Wilcoxon signed-rank testover the 235 test images, and the standardized calibration residual thatexposes a restorer whose restoration quality conceals a mismatch. Threeexternal arms follow, each with its own protocol: INRIA Person with ninefrozen detectors trained on clean INRIA (90 test frames, run both aslossless Portable Network Graphics and with the JPEG stage added), CrowdHumanwith 500 crowded validation frames and four COCO-pretrained detectors appliedzero-shot, and CityPersons with a fixed 200-frame subset and threeCOCO-pretrained detectors, run as a pre-specified go/no-go replication gate. Note on the arms and what they may be compared against. Two arms carry scopelimits that are stated in the article and repeated here so that no number inthis package is read outside them. The CityPersons arm uses a custom personAP@50 protocol on a 200-frame subset, chosen for parity with the other armsof this study; it is NOT the official CityPersons log-average miss-rate(MR^-2) protocol, and its absolute numbers are not comparable to thatbenchmark's leaderboard. Its pre-specified verdict was partial replication,reported as such: the matched-versus-blind separation reproduces in all ninedetector-by-sigma cells and the pooled gap grows monotonically, but thesigma = 30 gap of +0.072 AP@50 falls 0.008 short of the pre-stated +0.08threshold and PSNR does track the benefit there. The INRIA arm is a nullresult and is reported as one: the blind restorers are effectively matched onpure Gaussian INRIA, so the discriminator correctly does not separate them.That arm changes the dataset and the degradation recipe at the same time --the INRIA noisy sets were stored losslessly, without the JPEG re-encode ofthe PnPLO recipe -- so it cannot by itself attribute the vanishing separationto dataset easiness. The INRIA+JPEG cross-cell in this package(the two p14_inria_jpeg_* files) is what separates those two factors. Note on the image data: this package provides the quantitative results, theown-trained checkpoints and the source code, not the underlying images. PnPLOis a publicly released benchmark (Karthika and Chandran, 2020; distributed as"karthika95/pedestrian-detection"). INRIA Person, CrowdHuman and CityPersonsare third-party datasets available from their original providers, all citedin the article; CityPersons is built on Cityscapes frames and carries anon-commercial research license, so no frame of it is redistributed here. Thenoise is deterministic, but the seeding differs by set and only part of thegenerating code is in this package: (a) the core PnPLO grid -- thenoise_10/20/30 sets that p14_gate_table.csv, p14_sps_per_image.csv andp14_1c_corrected_detection.csv are computed on -- was produced by theparent-study notebook, which is not included, with a global numpy seed equalto sigma (np.random.seed(sigma)) and sequential draws in that notebook'sdataset order; (b) the arm sets generated inside the main pipeline notebook(its arms and gate cells) and the two INRIA arms use a per-image generatorseeded by the tuple (global seed 140704, sigma, the 32-bit cyclic redundancycheck of the image key), which can be regenerated image-for-image without thestored sets; (c) the CrowdHuman arm draws from numpy default_rng(0)sequentially over its frames (source_code/task_a_crowdhuman/prep_crowdhuman.py); (d) the CityPersons arm seeds per image with the CRC32of "stem|sigma" (source_code/task_a_citypersons/prep_citypersons.py); (e) thenoise-level estimator characterization uses default_rng(20260720 + sigma)(source_code/sigma_hat/validate_sigma.py). Note on third-party denoiser code and checkpoints: the deep restorers SwinIR,SCUNet, FFDNet, Restormer, PromptIR and NAFNet are run from the weightspublished by their original authors, which are not redistributed here. Theexact checkpoint file used for each of them, its training corruption and thelevel at which it was applied are recorded in the checkpoint-provenance tableof the article, and the download locations are given in the notebook headers.The architecture definitions of these restorers are likewise notredistributed: the notebooks fetch the model files at run time from theauthors' repositories on their default branch (Restormer, SwinIR, PromptIR,SCUNet, KAIR for FFDNet, and NAFNet), so a reader re-running a denoisingbundle after an upstream refactor may need to retrieve the model file fromthe repository revision current in July-August 2026; every load usesstrict=True, so a mismatch raises rather than silently substituting adifferent network.What this package does contain are the weights the core grid depends on andthat are not public releases: the three own-trained denoisers (DnCNN, CAE,CAE-PSO) at each noise level used in the article, and the frozen YOLOv9-Mdetector of the core grid. The eight further PnPLO detectors and the nineINRIA detectors of the cross-detector and cross-dataset arms are likewiseauthor-trained but are withheld for size and available on request (see thecheckpoints/detector_yolov9m_clean entry below). Note on the two evaluation pipelines: the three own-trained convolutionalbaselines run sliding-window inference, and an earlier revision of thattiling under-covered the bottom and right image margins. All fidelity anddetection numbers reported in the article are computed on outputsregenerated with the corrected tiling, and detection for the nine affectedcells was re-evaluated on those same corrected outputs in a singleultralytics 8.4.98 session, so that fidelity and detection are read from thesame frames in the same session. Both passes are kept in this package andare distinguished by a column, not by a folder: indata/pnplo/p14_fidelity_test235_fixed.csv the column "variant" takes thevalues old (the legacy tiling), fixed (the corrected tiling) and control (theclassical outputs, which never used the tiled path and reproduce their priorfidelity to within 0.001 dB), and indata/pnplo/p14_1c_corrected_detection.csv the column "arm" takes the valuesnoisy, old, fixed, control_old and panel. The experimental design isidentical across the two passes; only the input frames differ. One provenancecaveat is disclosed in the article and applies to these files: the correctedoutputs were regenerated on different hardware with mixed precision, so theresulting shifts at sigma = 30 exceed the isolated margin effect (about+0.009 mAP@50 measured same-hardware) and reflect regeneration provenance aswell as the tiling correction. Note on library versions: four ultralytics releases are recorded for thearms of record other than the core grid (8.4.90, 8.4.96, 8.4.98 and 8.4.105),and each of those arms is read within one release so that evaluator driftcancels (the detector checkpoint itself was trained under 8.4.58, and thesuperseded p14_qual_grid.py ran under 8.4.14); the core grid's own release isnot recorded, see below. The control re-evaluations andthe permutation gate were run on ultralytics 8.4.98 (PyTorch 2.11, CUDA12.8), on which the noisy baselines of the main grid reproduce to thereported precision (0.399 at sigma = 30), indicating no version drift in theevaluator. The core grid itself was produced by the main pipeline notebook,which installs ultralytics unpinned and does not record its version; itsYOLOv9-M cells are reproduced in-session by the cross-detector sweep (pinnedto 8.4.90; SwinIR / noisy 0.638 / 0.399 at sigma = 30) and by the 8.4.98re-evaluations. The two INRIAarms and the cross-detector sweep were run on ultralytics 8.4.90 (pinned insource_code/inria_cross_dataset, source_code/task_1b_inria_jpeg andsource_code/cross_detector). The CrowdHuman arm was run on ultralytics 8.4.96and the CityPersons gate arm on ultralytics 8.4.105, each verdict being amatched-minus-blind gap read within its one version; those two notebooksinstall ultralytics unpinned and print the installed version at setup, so theversions of record for them are the ones stated here and in the article'sexperimental-setup section, not recorded in their CSV files. Detectionevaluation in the core grid, the control re-evaluations, the permutationgate, the cross-detector sweep and the two INRIA arms uses the ultralyticsvalidator at conf = 0.25, IoU = 0.5, input size 640. The CrowdHuman andCityPersons arms are the exception: their detections are collected withmodel.predict at conf = 0.001, NMS IoU 0.7, input size 640, and scored by aself-contained PASCAL VOC person AP@50 (source_code/task_a_crowdhuman/eval_harddata.py and source_code/task_a_citypersons/eval_citypersons.py;column AP50_person in the two CSVs), so their absolute values are notcommensurable with the validator-derived mAP@50 of the core grid -- onlydirections and relative gaps may be compared across arms; the samepredict-and-VOC path also produced data/pnplo/p14_personlike_diag.csv (seeits entry). The file p14_1c_corrected_detection.csv carries the evaluator version per row, andresults/task_d_env.txt records the full interpreter, library, seed andcheckpoint set of the Task B/C/D runs. When using the source code, please follow the instructions provided in thenotebook header, and in the WORKFLOW.md file that accompanies each taskbundle under source_code -- those files give the run order and the smoke testfor their task. Those WORKFLOW.md files, the notebook headers and the taskprotocols under protocol/ are the contemporaneous internal work-orders, keptfor provenance; parts are in the authors' working shorthand, two of them(task_bcd, task_letterbox) were written in Vietnamese and are provided inEnglish translation, and the letterbox notebook's own header cell is left inthe original shorthand; the inline comments and progress-print messages ofthe main pipeline notebook (source_code/P14_SPH_Modern_Denoisers_Experiments.ipynb) are likewise in Vietnamese throughout, as are its markdown sectioncells -- the code itself is unaffected. "Barry" / "Barry-main" in them is the authors'internal name for their scripted-analysis tooling, not a person; aninstruction to "send the printout to Barry" means: stop, do not loosen thecheck, and contact the deposit's contact author at kiet.vo.thanh.st@vsb.cz.The standalone figure scripts in source_code/figure_scripts regenerate thecustom-coded figures of the article. Every input they need is present inthis package EXCEPT for the qualitative-figure scripts(source_code/task_qual_figs/qual_figs.py and the earlier p14_qual_fig.py andp14_qual_grid.py), which need the image sets -- those are third-partymaterial and are not redistributed here. The scripts alsocarry absolute paths from the authors' working tree and must be repointedbefore they run; see source_code/figure_scripts/00_PATHS_README.txt for themapping. The manifest folder holds a SHA-256 listing of every file togetherwith the script that produced it, so the integrity of the package can bechecked after download by running "python make_manifest.py --check" frominside that folder. IMPORTANT: Do not modify the file locations within the folder structure. Folder Structure-------------------------------------------------------------------------------- 14_sph_modern_denoisers_dataset (contains 7 subfolders, readme.txt and LICENSE.txt)│├── data (per-image and per-cell measurements, one subfolder per benchmark)│ ├── pnplo (core grid: fidelity, structure, detection, both pipelines)│ │ ├── p14_gate_table.csv, p14_sps_per_image.csv│ │ ├── arms_rows.csv, arms_latency.csv, arms_latency_meta.json│ │ ├── p14_1c_corrected_detection.csv, p14_fidelity_test235_fixed.csv│ │ ├── xdet_full9_paper.csv│ │ └── task_c_personlike_v2.csv, p14_personlike_diag.csv│ ├── inria (cross-dataset arm, lossless and JPEG cross-cell)│ │ ├── p14_inria_xdataset.csv, p14_inria_imgmetrics.csv│ │ └── p14_inria_jpeg_xdataset.csv, p14_inria_jpeg_imgmetrics.csv│ ├── crowdhuman│ │ └── p14_harddata_crowdhuman.csv│ └── citypersons│ ├── p14_citypersons.csv│ └── citypersons_provenance_manifest.json│├── protocol (the dated pre-specification of the gated arms)│ ├── citypersons_gate_protocol.md, crowdhuman_arm_protocol.md│ ├── crowdhuman_prediction_2026-07-09.md (the dated work-order)│ └── README.md (what they pre-state, and when)│├── results (aggregated statistics, gate verdicts and the run environment)│ ├── arms_gate_v2_verdict_LEGACY_precorrection.json│ │ (SUPERSEDED pre-correction run; the gate│ │ values of record are in step2c/step2d)│ ├── step2_results.json, step2c_1c_results.json, step2d_results.json│ ├── scunet_real_residuals.json│ ├── p14_fidelity_gates.json, README_p14_fidelity.md│ ├── perm_fixed_result.json, ssim_fixed_result.json, spsedge_fixed_result.json│ ├── task_b_permutation.json (SUPERSEDED sampled pre-correction draw)│ ├── task_d_env.txt│ ├── sigma_hat_characterization.txt│ ├── sigma_hat_frames_n120.csv, sigma_hat_summary_n120.json│ └── qual_provenance.json (audit record of the two qualitative figures)│├── checkpoints (the core-grid weights; see the checkpoints entries)│ ├── denoisers_own (DnCNN, CAE, CAE-PSO at sigma = 10/20/30)│ │ ├── dncnn_sigma{10,20,30}.pt, autoencoder_sigma{10,20,30}.pt│ │ └── cae_pso_sigma{10,20,30}.pt + matching _params.yaml│ └── detector_yolov9m_clean (the frozen detector of the core grid)│ ├── weights/best.pt│ └── args.yaml, results.csv│├── source_code│ ├── P14_SPH_Modern_Denoisers_Experiments.ipynb (main pipeline; cell outputs│ │ cleared -- its results ship in data/ and results/)│ ├── task_1b_inria_jpeg/, task_1c_corrected_detection/,│ │ task_a_crowdhuman/, task_a_citypersons/, task_bcd/, task_letterbox/,│ │ task_qual_figs/ (regenerated the two qualitative figures, 2026-08-29)│ │ (each: WORKFLOW.md + build_notebook.py + the Colab notebook + helpers)│ ├── cross_detector/, inria_cross_dataset/ (the two cross-arm notebooks)│ ├── stats/ (the aggregation scripts)│ ├── sigma_hat/ (the noise-level estimator)│ └── figure_scripts/ (standalone figure generators)│├── smoke (side outputs -- NOT article inputs; see the warning below)│├── manifest│ ├── MANIFEST.csv (path, bytes, SHA-256 for every file)│ └── make_manifest.py│├── LICENSE.txt└── readme.txt File Descriptions-------------------------------------------------------------------------------- • data/pnplo/p14_gate_table.csv The core grid of the study: 24 rows = 8 denoisers (SwinIR, Restormer, PromptIR, BM3D, Gaussian filter, DnCNN, CAE as "autoencoder", CAE-PSO) x 3 noise levels, each row aggregated over all 235 test images. Columns: cond, method, sigma, family (classical or modern), n_img, sps_grad, sps_edge, psnr, ssim, map50, f1, d_map50, d_f1. The two d_ columns are the within-sigma changes against the noisy baseline at the same level; they, not the raw map50, are the quantity the article analyses. This file is the raw source of the predictor-comparison and dose-response figures and of the appendix grid table. IMPORTANT -- NINE SUPERSEDED CELLS. For the three own-trained denoisers (autoencoder, dncnn, cae_pso) at sigma = 10/20/30, the d_map50 and d_f1 columns of THIS file are the legacy tiled-inference pass and are NOT the values the article reports. The article uses the corrected-output pass: see data/pnplo/p14_1c_corrected_detection.csv and results/step2c_1c_results.json. The difference is material and is the subject of an explicit retraction in the article: this file gives dncnn at sigma = 30 as d_map50 = +0.149161, whereas the article reports -0.054 -- which is why the article says degradation matching "is the stronger determinant of" benefit and "does not by itself decide" it. Do not quote the nine own-trained d_ cells of this file. • data/pnplo/p14_sps_per_image.csv Per-image fidelity and structure measurements behind the aggregate rows above: 5,640 rows, columns cond, image, sps_grad, sps_edge, psnr, ssim. • data/pnplo/arms_rows.csv The auxiliary arms of the core grid: 48 rows, columns cond, arm, sigma, map50, f1, psnr, sps_grad, n_img, variant, model, sigma_w, role, mismatch. "arm" identifies the experiment (arm1 and arm1_base = the controlled JPEG ablation and its baseline; arm2, arm3, arm3b = the recipe-crossing arms; e1_ffdnet, e2_nafnet = the two added restorers; e3_qsweep = the JPEG quality sweep at q = 75 and 85; e4_sigma50 = the sigma = 50 stress point). "variant" is jpeg or pure on the arm1 / arm1_base rows (the ablation lever itself) and is empty on every other arm; sigma = -1 denotes a clean reference row; "sigma_w" is the level the restorer's weights were trained for and "mismatch" the resulting gap, so that degradation matching can be read directly off the file. • data/pnplo/arms_latency.csv, arms_latency_meta.json Measured per-image runtime of eight restorers (BM3D, FFDNet, Gaussian filter, NAFNet-SIDD, PromptIR, Restormer, SCUNet-25, SwinIR-25). The first column is unnamed and holds the restorer name; the remaining columns are mean_ms, median_ms, std_ms and n (the number of timed images). These feed the cost-versus-recovery figure and the deployment table; the metadata file records the timing conditions. • data/pnplo/p14_1c_corrected_detection.csv Detection on the corrected outputs: 33 evaluation cells, columns dataset, detector, arm, method, sigma, map50, map50_95, n_images, ultralytics, ts. All rows are PnPLO-test235 under yolov9m_clean at ultralytics 8.4.98. The "arm" column separates the two pipelines (see the note above): noisy, old, fixed, control_old, panel. • data/pnplo/p14_fidelity_test235_fixed.csv Fidelity of the regenerated outputs: 24 rows, columns method, sigma, variant, psnr_mean, sps_mean, n_images, with variant in {old, fixed, control}. Read together with the file above, this is what allows fidelity and detection to be quoted from the same frames. • data/pnplo/xdet_full9_paper.csv The cross-detector arm: 108 rows = 9 frozen PnPLO detectors (YOLOv8m, YOLOv9m, YOLOv10m, YOLO11m, YOLO12m, YOLO26m and three ResNet-backbone variants) x 3 noise levels x 4 conditions (noisy, matched_swinir, blind_restormer, blind_promptir). Columns: detector, sigma, condition, map50. This file is the sole source of the cross-detector section and its figure. • data/pnplo/task_c_personlike_v2.csv, p14_personlike_diag.csv The class-agnostic control: person-only versus merged person / person-like scoring, used to show that the two-class label space of PnPLO is not what drives the matched-versus-blind gap. The two files come from two different evaluators and are not directly comparable: task_c_personlike_v2.csv (columns detector, cond, mode, ap50, map50_all) is the ultralytics validator at conf = 0.25 run by task_bcd; p14_personlike_diag.csv (columns detector, cond, AP_person, AP_merged, conf_person_to_personlike) is the earlier diagnostic, not quoted in the article, kept because it additionally records the person-to-person-like confusion rate; it was produced by diag_personlike.py -- detections collected with model.predict at conf = 0.001, NMS IoU 0.7 and scored by the same self-contained VOC AP@50 that the CrowdHuman and CityPersons evaluators later copied -- and that script is not included. The suffix _v2 marks the revised run that the article uses; the superseded first run is not included. • data/inria/p14_inria_xdataset.csv, p14_inria_imgmetrics.csv The INRIA arm, stored losslessly as PNG: 117 detection rows = 9 detectors x 13 conditions (columns detector, condition, sigma, map50, f1; condition in {clean, noisy, swinir, restormer, promptir}; sigma = -1 marks the clean row), and 12 image-metric rows over the 90 INRIA test frames (columns condition, sigma, n, psnr, ssim, sps_grad). This is the null arm. • data/inria/p14_inria_jpeg_xdataset.csv, p14_inria_jpeg_imgmetrics.csv The same INRIA arm with the JPEG acquisition stage added, in the same schema and the same row counts. Comparing this pair against the pair above is what separates "different dataset" from "different degradation recipe". • data/crowdhuman/p14_harddata_crowdhuman.csv The CrowdHuman replication: 72 rows = 4 COCO-pretrained detectors x 3 noise levels x 6 conditions (noisy plus five restorers), over 500 crowded validation frames. Columns: dataset, detector, sigma, condition, cls (matched, blind, or "-" for the noisy baseline), AP50_person, ddet, SPSgrad, SPSedge, PSNR, SSIM. • data/citypersons/p14_citypersons.csv The pre-specified replication gate on driving scenes: 54 rows = 3 COCO-pretrained detectors x 3 noise levels x 6 conditions, over the fixed 200-frame subset, in the same schema as the CrowdHuman file plus an n_images column. Read only with the scope limit stated above: this is a custom person AP@50 protocol, not the official CityPersons MR^-2 protocol. • data/citypersons/citypersons_provenance_manifest.json The provenance guard written when that subset was prepared: the annotation source, the class and size filters, and the run configuration. (Renamed from manifest.json on packaging, to keep it distinct from the package manifest in the manifest folder; its contents are unchanged.) • protocol/ The dated pre-specification of the two gated external arms, deposited as the article states. citypersons_gate_protocol.md (header dated 2026-07-20, annotation-source verification 2026-07-24) fixes the CityPersons go/no-go gate -- the "matched-minus-blind gap > +0.08 and grows" criterion and its REPLICATES / Partial / FAILS reading -- before any CityPersons data existed; crowdhuman_arm_protocol.md is the operational protocol of the CrowdHuman arm (roster, detectors, recipe, smoke-first and resume discipline); and crowdhuman_prediction_2026-07-09.md is the dated work-order, written five days after the V1-V4 gate specification (locked and executed on 2026-07-04, pre-registration commit be958c5, as recorded in source_code/P14_SPH_Modern_Denoisers_Experiments.ipynb), that pre-states the expectation the CrowdHuman subsection quotes ("expect matched d_mAP50 > +0.10 while blind d_mAP50 is about 0 at sigma = 30"). README.md in that folder says what each file pre-states and when, and gives the timeline corroboration. All three are reproduced as written, including their operational notes in the authors' working shorthand. The first two are verbatim copies of the WORKFLOW.md files in source_code/task_a_citypersons/ and source_code/task_a_crowdhuman/ (identical checksums in the manifest), placed here so the pre-specification is findable. • results/arms_gate_v2_verdict_LEGACY_precorrection.json The V1 to V4 criteria as they stood BEFORE the corrected-output detection pass of Section 4 replaced the detection entries of the three own-trained learned baselines. It is SUPERSEDED for every detection-side entry and is NOT the file behind the gate-summary table, with one exception noted below: its V1 rho (0.7857), its permutation p (1e-4) and its V4 verdict (PASS) all differ from the published 0.821, 3.1e-5 (within-stratum) / 3.4e-3 (blocked-by-denoiser) and "not resolved". It is retained as a record of the earlier run and of what the correction changed. The gate values of record are in step2c_1c_results.json and step2d_results.json. Exception: the V3 fidelity Wilcoxon p (1.3228e-40 for Restormer and for PromptIR) is the value printed in the gate-summary table, and this file is its file of record. It is a paired test on the per-image PSNR of the Restormer and PromptIR JPEG-ablation arms, which the corrected-output pass did not touch (that pass replaced only the detection entries of the three own-trained baselines), so the scalar is unchanged. The 235 per-image PSNR pairs behind it are computed in memory by source_code/P14_SPH_Modern_Denoisers_Experiments.ipynb (cell 14, _per_image_psnr, on the a1_*_pure_30 and a1_*_jpeg_30 arm outputs) and were not persisted; regenerating them needs the third-party Restormer and PromptIR weights and the image corpus, neither of which is redistributed. • results/step2_results.json, step2c_1c_results.json, step2d_results.json The aggregated statistics computed by the three scripts in source_code/stats, keyed by the quantity they produce: the CityPersons pooled and per-detector gaps, separation counts, blocked and sign tests, stratified and per-metric correlations (step2); the integration of the corrected-output pass, including the old-versus-new value of every affected number and the session-equivalence check (step2c); and the discriminator table together with the V2, V4 and NAFNet replacements (step2d). Two keys of step2_results.json are pre-correction and are NOT the values the article reports: pnplo_v1_blocked (k = 4 of 5040, p = 7.9e-4) and pnplo_blocked_restorer (T = +0.076, k = 1 of 36, p = 0.028) were computed before the corrected-output pass -- step2_stats.py patches sps_grad from the fixed fidelity CSV but leaves the nine own-trained d_map50 cells at their legacy values. They are superseded by step2c_1c_results.json, which stores them as v1_old / blocked9_old beside the values of record v1_new (k = 17 of 5040, p = 3.4e-3) and blocked9_new (T = +0.058, k = 8 of 36, p = 0.22, reported in the article as not significant). All city_* and crowdhuman_blocked keys of step2_results.json are current. • results/scunet_real_residuals.json The residuals of the real-noise SCUNet arm and the exact permutation tail quoted in the gate section. • results/p14_fidelity_gates.json, README_p14_fidelity.md The acceptance gates of the tiling regeneration, and a short written summary of the regenerated fidelity table. • results/ssim_fixed_result.json, spsedge_fixed_result.json The SSIM and SPS_edge correlation results recomputed on the corrected-tiling outputs, per metric. • results/perm_fixed_result.json The stratified permutation as it stood when only the FIDELITY side had been regenerated: its FIXED block reads SPS_grad from the corrected outputs, but d_map50 is still the legacy tiled-inference pass. Its rho (0.8214) coincides with the published value, but its per-sigma profile (0.75 / 0.75 / 0.964) and its p (6.0e-5) are NOT the published gate and must not be quoted as such. It is retained as a record of what the detection correction changed. The V1 values of record are rho_strat = 0.821 with per-sigma 0.857 / 0.750 / 0.857 (step2c_1c_results.json, "v1_new"); the like-for-like within-stratum exact p is 3.1e-5 (scunet_real_residuals.json, "exact_perm_note") and the separate blocked-by-denoiser p is 3.4e-3 ("v1_new", k = 17 of 5040); the same file's "v1_old" reproduces this file's rho and per-sigma. Its OLD block is the sampled pre-correction draw described next. • results/task_b_permutation.json A SAMPLED within-stratum permutation (NPERM = 100,000, seed 2026) of the PRE-CORRECTION gate: rho_obs = 0.7857, per-sigma 0.750 / 0.750 / 0.857, k_tail = 10, p = 1.1e-4. SUPERSEDED, and NOT the tail behind the gate section: the article's V1 is rho_strat = 0.821 (per-sigma 0.857 / 0.750 / 0.857) with an exactly enumerated p = 3.1e-5. The rho and per-sigma of record are in step2c_1c_results.json ("v1_new") and the exactly enumerated tail is in scunet_real_residuals.json ("exact_perm_note", P(>= obs) = 3928502 / 128024064000 = 3.07e-5). This file is the OLD block of perm_fixed_result.json, retained as a record of the earlier run. (Its internal key is named "p_exact" although the computation is sampled.) • results/task_d_env.txt The recorded run environment: interpreter, numpy, pandas, scipy, ultralytics, torch, CUDA and GPU, the evaluation protocol, the noise seed and the per-image generator construction, the permutation seed, and the detector checkpoint paths. Its noise-seed line is a constant transcribed by the task_bcd notebook and describes construction (b) of the note on the image data (the arm sets and the INRIA arms); the sigma = 30 core-grid condition folders that Tasks B and C read were generated by the parent-study pipeline under construction (a). • results/sigma_hat_characterization.txt The characterization of the noise-level estimator: the clean-image baseline, the clipping diagnostic, the raw estimate per level under both storage recipes, and the calibrated routing accuracy. It is the artifact behind the two routing figures quoted in the method section (88.5 percent of frames routed to the exact grid level and 99.4 percent to within one level, over 120 frames x 4 levels = 480 decisions). Its header records the frame selection, the seed and the exact command that produced it, so the file can be regenerated with source_code/sigma_hat/validate_sigma.py. • results/sigma_hat_frames_n120.csv, sigma_hat_summary_n120.json The per-frame evidence behind the same two figures, so they can be checked without re-running anything. The CSV has 1440 rows = 120 frames x 4 noise levels x 3 (calibration recipe, storage recipe) blocks, with the estimate, the level it routed to, and whether that was exactly or adjacently correct. Group by the PAIR (cal_store, store), not by store alone: jpeg95/jpeg95 gives 425/480 = 88.5 percent exact and 477/480 = 99.4 percent adjacent, which are the article's numbers; lossless/lossless gives 91.0 and 100.0; and the deliberately mismatched lossless/jpeg95 block gives 29.0 and 95.4, which is the finding that the calibration must be built on the deployment storage recipe. The summary JSON carries the per-level means and the midpoint thresholds. • results/qual_provenance.json The audit record of the 2026-08-29 run that regenerated the two qualitative figures (the single-frame comparison, frame 45, and the three-row grid, frames 100 / 81 / 62). For each displayed frame it lists every candidate panel source with its PSNR against the clean frame and against the grid noisy input, and the number of person detections and their confidences from the frozen detector, so the panel actually used can be checked against the alternatives that were rejected. It also records the panel source directories, the output file sizes, the drawn and printed type sizes with the 6 pt floor check, the resolved font and the ultralytics version (8.4.98) of the regenerating run. See source_code/task_qual_figs/WORKFLOW.md for what that run changed and why. • checkpoints/denoisers_own/ The three own-trained denoisers at each noise level used in the article: DnCNN (17 layers, 64 features, residual), the convolutional autoencoder, and CAE-PSO, whose selected architecture and learning rate are recorded in the matching _params.yaml file. All were trained on the PnPLO training split on random 128x128 patches with on-the-fly Gaussian noise at the target level, Adam (learning rate 1e-3 for DnCNN and the autoencoder; for CAE-PSO the PSO-selected rate recorded in its _params.yaml, 3.96e-4 / 8.22e-4 / 5.54e-4 at sigma = 10 / 20 / 30; reduce-on-plateau schedule with patience 5 and factor 0.5), mean-squared-error loss, 50 epochs, batch size 256 -- the parent-study notebook selects the batch size from GPU memory, and the run that produced these checkpoints used an NVIDIA A100 80 GB. That notebook is not included here; the checkpoints and their _params.yaml files are. (The notebook's training functions carry defaults of 30 epochs and batch size 16, which the training cell overrides.) Checkpoints at sigma = 1 and sigma = 5 exist from the parent study but are not used anywhere in this article and are therefore not included. • checkpoints/detector_yolov9m_clean/ The frozen YOLOv9-M detector of the core grid, trained on clean PnPLO and never fine-tuned under any denoiser: weights/best.pt (about 41 MB) together with the exact training arguments (args.yaml) and the training log (results.csv). Without this file no d_mAP50 of the core grid can be reproduced. The eight other PnPLO detectors and the nine INRIA detectors of the cross-detector and cross-dataset arms are not included here for size reasons and are available from the corresponding authors upon reasonable request; the CrowdHuman and CityPersons arms use COCO-pretrained detectors that ultralytics downloads automatically. • source_code/P14_SPH_Modern_Denoisers_Experiments.ipynb The main pipeline notebook of this study: the setup and XAI cells copied from the parent-study notebook, the modern-denoiser backend (SwinIR, Restormer, PromptIR, with the SCUNet, FFDNet and NAFNet arms), the per-image fidelity and structure measurements, the gate and the arms, producing p14_gate_table.csv, p14_sps_per_image.csv and the arms files. It READS the core-grid noisy sets and the five classical / own-trained denoised sets (Gaussian filter, BM3D, autoencoder, DnCNN, CAE-PSO) from disk: those sets, the training of the three own-trained denoisers and the core-grid noise generation were produced by the parent-study notebook, which is not included in this package. It predates the task bundles below, which were added later and run as separate notebooks. • source_code/task_*/ One self-contained bundle per follow-up experiment, each with its own WORKFLOW.md giving the run order and the smoke test, a build_notebook.py that generates the Colab notebook, the notebook itself, and any preparation or evaluation helper: task_1b_inria_jpeg (the INRIA+JPEG cross-cell), task_1c_corrected_detection (detection on the corrected outputs), task_a_crowdhuman and task_a_citypersons (the two zero-shot replication arms, each with its dataset preparation script), task_bcd (the sampled pre-correction permutation, the person-like control task_c_personlike_v2.csv and the environment record), task_letterbox (the regeneration of the fidelity measurements with the corrected tiling), and task_qual_figs (the 2026-08-29 regeneration of the two qualitative figures; its WORKFLOW.md records what changed and why, and it needs the image sets, which are not redistributed). The outputs of these bundles are not duplicated inside the bundles; they have been promoted into data/ and results/ as described above. • source_code/cross_detector/, source_code/inria_cross_dataset/ The notebooks that produced xdet_full9_paper.csv and the two INRIA files, with their builders and the custom ResNet module definitions that are required to load the three ResNet-backbone detectors. Every builder in this package writes its notebook beside itself; build_inria_colab.py takes its ResNet-loader cell from ../cross_detector/P14_CrossDetector_Colab.ipynb (cell 2), so the folder pair is self-contained. The notebooks themselves are Colab notebooks and read the image sets and detector weights from Drive paths, which are not part of this package. • source_code/stats/ step2_stats.py, step2c_1c_integration.py and step2d_full_replacements.py, the aggregation scripts that produced results/step2_results.json, step2c_1c_results.json and step2d_results.json. A byte-identical copy of step2d_results.json also sits in this folder beside step2d_full_replacements.py (same SHA-256 in the manifest); results/step2d_results.json is the copy of record, and the other two scripts' outputs ship only under results/. Like the figure scripts they carry absolute paths from the authors' working tree rather than reading data/ relatively, so they must be repointed before they run; the inputs they need are the corresponding files under data/ and results/. • source_code/sigma_hat/ sigma_estimator.py implements the median-absolute-deviation estimator on the diagonal detail subband of a Haar wavelet transform; validate_sigma.py characterizes it and produces results/sigma_hat_characterization.txt. Set IMG_DIR at the top of validate_sigma.py to a directory holding the PnPLO test frames before running it. • source_code/figure_scripts/ Standalone Python scripts (matplotlib) that regenerate the custom-coded figures of the article. They carry absolute paths from the authors' working tree and must be repointed to this package before they run -- see 00_PATHS_README.txt in that folder for the mapping. The scripts are: make_p14_figs_v2.py (the predictor-comparison, recipe-crossing, cost, JPEG-sweep and sigma = 50 figures), p14_stats_fig2.py (the dose-response figure), make_p14_xdet_fig_v2.py and make_p14_inria_fig_v2.py (the cross-detector and cross-dataset figures, in the version that prints the numeric value above each bar), make_p14_crowdhuman_fig.py, and p14_qual_fig.py and p14_qual_grid.py (the earlier versions of the two qualitative figures; the figures in the article were regenerated on 2026-08-29 by source_code/task_qual_figs/qual_figs.py, and both generations need the image sets). qual_scan_cache.csv beside them is a side output of p14_qual_grid.py, not an article input: the script writes it on its first pass over the 235 PnPLO test frames and reads it on later runs to skip the rescan; its columns are n (frame number), conf (highest person confidence on the clean panel) and gtfrac (largest ground-truth box as a fraction of the frame), and it is read from the script's own folder, so it needs no repointing. p14_fixed_patch.py is the helper that make_p14_figs_v2.py and p14_stats_fig2.py import to patch the corrected-tiling fidelity values into the gate table. make_p14_f8.py draws the mechanism figure of the conclusion (fig_mechanism.pdf). ieee_style.py and p14_barlabels.py hold the shared plot style and the bar labelling. Two classes of figure are not reproducible from this package alone: the qualitative ones need the image sets, and the schematic figures are drawn in the manuscript source rather than by a script. • smoke/ WARNING: no table or figure of the article is derived from any file in this folder. Six of the eight files are genuine smoke-test outputs -- the tiny end-to-end validation pass a notebook runs before the real job, covering a handful of images and one or two conditions, with numbers that are meaningless as results. Two are not, and are kept here only because they are side outputs rather than article inputs: - p14_1c_corrected_detection_SMOKE.csv is the task-1c smoke pass, whose smoke lever is the cell count (3 of 33), not the image count: every row is the full 235-image test split, and its three map50 values reproduce exactly the noisy@30, fixed dncnn@30 and old dncnn@30 rows of data/pnplo/p14_1c_corrected_detection.csv. Quote that file, not this. - task_b_permutation_LOCAL.json is a local full-N replication of the sampled pre-correction permutation (NPERM = 100,000) whose p matches results/task_b_permutation.json; the genuine smoke output of that task is task_bcd_smoke_task_b_permutation.json (NPERM = 1,000). They are included only so that a reader re-running a notebook can compare their own pass against ours. Every file here is marked _SMOKE, _smoke or _LOCAL in its name; no file outside this folder is one. • manifest/MANIFEST.csv, manifest/make_manifest.py A listing of every file in the package with its size in bytes and SHA-256 digest, and the script that generates it. Run "python make_manifest.py" from inside the manifest folder to regenerate the listing, or "python make_manifest.py --check" to verify a downloaded copy against it; the check reports any missing, extra or altered file and exits non-zero if the package does not match. Article cross-reference-------------------------------------------------------------------------------- Figure and table numbers are not used below, because they shift with thejournal template; each item is named by its subject instead. • Predictor comparison (within-sigma correlation of SPS_grad, SSIM and PSNR with d_mAP50) <- data/pnplo/p14_gate_table.csv, results/step2c_1c_results.json• Dose-response of d_mAP50 <- data/pnplo/p14_gate_table.csv, data/pnplo/p14_fidelity_test235_fixed.csv, results/step2d_results.json, results/ssim_fixed_result.json, results/spsedge_fixed_result.json• Controlled JPEG ablation and the recipe-crossing arms <- data/pnplo/arms_rows.csv• Cost versus recovery <- data/pnplo/arms_rows.csv, arms_latency.csv• JPEG quality sweep and the sigma = 50 stress point <- data/pnplo/arms_rows.csv• Cross-detector generalization <- data/pnplo/xdet_full9_paper.csv• Cross-dataset generalization (INRIA) <- data/inria/p14_inria_xdataset.csv and p14_inria_imgmetrics.csv; the JPEG cross-cell from the two p14_inria_jpeg_* files• CrowdHuman replication <- data/crowdhuman/p14_harddata_crowdhuman.csv• CityPersons gate <- data/citypersons/p14_citypersons.csv, results/step2_results.json• Discriminator table (standardized calibration residuals) <- results/step2d_results.json• Gate summary table (V1-V4) <- results/step2c_1c_results.json, results/step2d_results.json, results/scunet_real_residuals.json, results/arms_gate_v2_verdict_LEGACY_precorrection.json (V3 Wilcoxon p row only; see the note on that file above), data/pnplo/arms_rows.csv (the V3 difference-in-differences row)• Deployment table (latency and recovery) <- data/pnplo/arms_latency.csv, arms_rows.csv and p14_gate_table.csv• Appendix grid table <- data/pnplo/p14_gate_table.csv patched with data/pnplo/p14_fidelity_test235_fixed.csv, results/ssim_fixed_result.json, results/spsedge_fixed_result.json and results/step2c_1c_results.json (see source_code/figure_scripts/p14_fixed_patch.py). All four inputs are needed: the fidelity CSV carries only psnr/sps and has no d_map50 column, so patching from it alone leaves the nine superseded own-trained detection cells in place -- including dncnn at sigma = 30, which would read +0.149 instead of the published -0.054.• Noise-level estimator (routing accuracy) <- results/sigma_hat_characterization.txt• Person-like class-agnostic control <- data/pnplo/task_c_personlike_v2.csv• Qualitative panels (single frame and grid) <- results/qual_provenance.json (the image sets themselves are not redistributed)• The checkpoint-provenance table, the positioning table and the schematic framework figure are documentary and have no data file.-------------------------------------------------------------------------------- 7. Funding: This work was supported in part by the European Regional Development Fund under the project Research Platform for Digital Transformation and Society 5.0 CZ.02.01.01/00/23_021/0012599 within the Jan Amos Komensky Operational Program. This work was supported in part by the Ministry of Education of the Czech Republic (Project No. SP2026/012, SP2026/075).8. Licence and reuse This deposit is NOT uniform. LICENSE.txt in the archive root is authoritative; the summary is: * CC BY 4.0 -- data/, results/, smoke/, manifest/, readme.txt, and checkpoints/denoisers_own/ (the nine own-trained DnCNN / autoencoder / CAE-PSO checkpoints, plain PyTorch, verified free of any ultralytics or AGPL string). This matches the two earlier Zenodo records from this group, 10.5281/zenodo.20842362 and 10.5281/zenodo.20842879. * AGPL-3.0-or-later -- checkpoints/detector_yolov9m_clean/ and every script under source_code/ that imports ultralytics (enumerate with `grep -rl ultralytics source_code/`; 34 paths as of 2026-09-02, of which the .md/.txt hits and two .py files that only name the pinned version without importing the library (figure_scripts/make_p14_xdet_fig_v2.py, stats/step2c_1c_integration.py) are not covered work and stay CC BY -- see LICENSE.txt). This is not a choice we made: the detector checkpoint was produced with Ultralytics YOLO 8.4.58 and its own serialised metadata reads `license: AGPL-3.0 (https://ultralytics.com/license)`. * Not redistributed, therefore not licensed by us -- all image data. The four evaluation corpora (PnPLO, INRIA Person, CrowdHuman, CityPersons) and the third-party restorer checkpoints (SwinIR, SCUNet, Restormer, PromptIR, FFDNet, NAFNet) must be obtained from their original providers under their own terms. CityPersons carries a non-commercial research licence. If a single uniform licence is required, the way to get one is to remove checkpoints/detector_yolov9m_clean/ and the ultralytics-importing scripts, at the cost of no longer shipping the frozen detector the whole study depends on.

提供机构:
Zenodo
创建时间:
2026-09-09
二维码
社区交流群
二维码
科研交流群
商业服务