遇见数据集

Statistical Infrastructure and Quantitative Analysis of SSIM BlindSpots in 82M Block Database

收藏
Zenodo2026-03-16 更新2026-05-26 收录
官方服务:

资源简介:

AbstractThis report establishes the statistical foundation for a diagnostic system analyzing approximately 82 million samples per processing path, enabling a transition from traditional scalar metrics to multidimensional structural analysis. This dataset supports the Merutan Theory framework, enabling structural diagnostics such as Axis‑based decomposition, Tier‑3 state transitions, SSIM BlindSpot detection, and Silent Collapse quantification. Building on the v1.0 theoretical proof of SSIM’s inherent limitations, this release provides the empirical basis required for large‑scale structural evaluation. Data Scale and ConfigurationThe analysis evaluates four distinct AI models, with each model processed through two separate paths: Notta (Standard) and TTA (Test Time Augmentation).- Per Path: Approx. 82,000,000 blocks- Per Model: Approx. 164,000,000 total samples (Notta + TTA)- Aggregate Study Total: Approximately 656,000,000 data points analyzed across all models Technical Foundation and Analysis MethodologyA high‑performance DuckDB backend enables efficient aggregation and analysis of this large‑scale dataset. This infrastructure allows quantification of internal failure modes—SSIM BlindSpots and Silent Collapses—that remain invisible to conventional evaluation metrics. These analyses form the empirical basis for Axis1–Axis6 structural metrics and Tier‑3 state classification within the Merutan Theory framework. Raw Data Provenance and Reproducibility (v1.1 Update)This project conducts a comprehensive analysis across four different AI models, each evaluated under two processing paths—Notta (Standard) and TTA (Test Time Augmentation)—resulting in eight total configurations. Across these configurations, the system processes approximately 82 million blocks per path, yielding roughly 650 million data points overall. To protect unpublished findings currently being prepared for an academic paper, the raw CSV data files themselves remain unpublished at this stage. These files will be released in this repository following the publication of the associated research paper. In this version (v1.1), a complete SHA256 hash list of all raw output files (raw_csv_sha256.zip) is provided to establish priority and mathematically guarantee data integrity. These hashes uniquely identify the foundational AI output data prior to any statistical processing, ensuring that the raw data released in the future will be exactly identical to the data used in this research, regardless of DuckDB versions, import settings, or other software configurations. Furthermore, the source code for the engine used to generate this dataset is already publicly available through the related Zenodo record (17677441) and on GitHub, ensuring full transparency and reproducibility of the data generation process. Note on TerminologyIn this dataset and documentation, the term “Original” is used interchangeably with “Ground Truth (GT)”, referring to the high‑quality source images used as the baseline for analysis. Dataset and Mandatory Citation:- Source: NIH ChestX-ray8 (Hospital-scale chest x-ray database)- Citation: Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., & Summers, R. M. (2017). "ChestX-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thoracic diseases." Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3462–3471. - Download: https://nihcc.app.box.com/v/ChestXray-NIHCC Software and Tools UsedThe dataset was generated and processed using the following specialized tools and AI models:- Source Data Generation Tool — Custom application developed by the author for block‑level diagnostic CSV generationhttps://zenodo.org/records/17677441- Upscayl — Primary GUI/engine for image upscalingLicense: GNU AGPLv3- Real‑ESRGAN / SwinIR — Underlying AI architectures used for image restoration and enhancementLicenses: Real‑ESRGAN (Apache License 2.0) / SwinIR (Apache License 2.0) Credits & Citations- Data Generation Tool — Developed by the author- Upscayl — Developed by Nayam Amarshe and TGS963- SwinIR — Liang, J., et al. “SwinIR: Image Restoration Using Swin Transformer.” arXiv:2108.10257 (2021)- Real‑ESRGAN — Wang, X., et al. “Real‑ESRGAN: Training Real‑World Blind Iterative Image Restoration.” ICCV Workshops, 2021(https://github.com/xinntao/Real-ESRGAN) Contact : s.shiny.n.works@gmail.com 概要本レポートは、処理パスごとに約 8,200 万サンプルを解析する診断システムの統計的基盤を確立し、従来のスカラ値ベースの評価から、多次元的な構造解析へと移行するための基礎を提供します。本データセットは Merutan Theory(メルタン理論)を支えるものであり、Axis に基づく構造分解、Tier‑3 状態遷移、SSIM BlindSpot の検出、Silent Collapse の定量化などの構造診断を可能にします。また、v1.0 において示された SSIM の本質的限界の理論的証明を踏まえ、本リリースでは大規模構造解析に必要な実証データ基盤を提供します。 データ規模と構成本解析では 4 種類の AI モデルを対象とし、各モデルを Notta(標準処理)と TTA(Test Time Augmentation)の 2 つの処理パスで評価しています。- 各パスあたり:約 8,200 万ブロック- 1 モデルあたり:約 1 億 6,400 万サンプル(Notta + TTA)- 全体合計:約 6 億 5,600 万データポイントを解析 技術基盤と解析手法高性能な DuckDB バックエンドを用いることで、この大規模データセットの高速な集計・解析を実現しました。この基盤により、従来の評価指標では検出できなかった内部的な失敗モード――SSIM BlindSpot および Silent Collapse――を定量化することが可能となりました。これらの解析は、Merutan Theory における Axis1〜Axis6 の構造指標および Tier‑3 状態分類の実証的基盤を形成しています。 生データの来歴と再現性(v1.1 更新) 本プロジェクトでは、4つの異なるAIモデルに対し、それぞれ「Notta」および「TTA」を適用した計8パターンの網羅的解析(計8,200万ブロック、約6.5億データポイント)を実施しています。 現在執筆中の論文における未発表の知見を保護するため、これらの生データ(CSV形式)本体は、現時点では非公開とし、関連論文の発表に合わせて本リポジトリにて公開する予定です。 本バージョン(v1.1)では、先行権の確定とデータの真正性を数学的に保証するため、すべての生出力ファイルに対する SHA256 ハッシュリスト(raw_csv_sha256.zip) を先行して提供します。これらのハッシュは、アプリケーション側で統計処理が行われる前の基礎となる AI 出力データを一意に特定するものであり、将来公開される生データが本リサーチの内容と完全に同一であることを保証します。 なお、本データセットを生成したエンジンのソースコードは、すでに関連レコード(17677441)および GitHub にて公開されており、データ生成プロセスの透明性と再現性が担保されています。 用語に関する注記本データセットおよびドキュメントにおいて、“Original(オリジナル)” は “Ground Truth(GT / 正解画像)” と同義であり、解析の基準となる高品質なソース画像を指します。 データセットおよび必須引用文献 本研究の大規模構造診断には、以下のデータセットを使用しています。 出典: NIH ChestX-ray8 (Hospital-scale chest x-ray database) 引用文献: Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., & Summers, R. M. (2017). "ChestX-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thoracic diseases." Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3462–3471. ダウンロード: https://nihcc.app.box.com/v/ChestXray-NIHCC 使用ソフトウェアおよびツールデータセットの生成および処理には、以下の専用ツールおよび AI モデルを使用しました。- ソースデータ生成ツール:著者が開発したブロック単位診断 CSV 生成アプリhttps://zenodo.org/records/17677441- Upscayl:画像アップスケーリングの主要 GUI / エンジン- Real‑ESRGAN / SwinIR:画像復元・強調に使用された基盤 AI アーキテクチャ- ライセンス:Upscayl(GNU AGPLv3)/ SwinIR(Apache License 2.0) クレジットおよび引用- データ生成ツール:著者による開発- Upscayl:Nayam Amarshe および TGS963 による開発- SwinIR:Liang, J., et al. “SwinIR: Image Restoration Using Swin Transformer.” arXiv:2108.10257 (2021)- Real‑ESRGAN:Wang, X., et al. “Real‑ESRGAN: Training Real‑World Blind Iterative Image Restoration.” ICCV Workshops, 2021 (https://github.com/xinntao/Real-ESRGAN) Technical Report v1.1: Statistical Infrastructure and Quantitative Analysis of SSIM BlindSpots in 82M Block Database Contact : s.shiny.n.works@gmail.com

提供机构:
Zenodo
创建时间:
2026-03-06
二维码
社区交流群
二维码
科研交流群
商业服务