Dege-Kangyur Dataset: A High-Fidelity Layout-Aligned Image Dataset for Historical Tibetan Document Restoration
收藏资源简介:
The Dege-Kangyur dataset is a high-fidelity, layout-aligned image dataset designed for weak-supervised historical Tibetan document restoration, reconstruction, and generation tasks. It contains 500 high-resolution color image pairs (original degraded scans and reconstructed clean versions) derived from the Derge woodblock-printed edition of the Tibetan Buddhist Canon (Kangyur), printed on Stellera chamaejasme paper. All images preserve the original non‑rectangular page geometry (curvature, red bounding frames, physical edges) without any geometric rectification. The clean images are generated by a dedicated reconstruction pipeline that removes degradation factors (water stains, ink bleed‑through, fiber noise) while retaining authentic paper texture and layout. This Zenodo record contains seven ZIP archives: 1. `degraded1.zip` – First 250 full‑resolution degraded images (fixed resolution ~5000×1000 pixels, 250 png files)2. `degraded2.zip` – Remaining 250 full‑resolution degraded images (fixed resolution ~5000×1000 pixels, 250 png files)3. `clean1.zip` – First 250 full‑resolution clean reconstructed images, paired with degraded1.zip (average resolution ~1835×350 pixels, varies slightly by page, 250 png files)4. `clean2.zip` – Remaining 250 full‑resolution clean reconstructed images, paired with degraded2.zip (average resolution ~1835×350 pixels, varies slightly by page, 250 png files)5. `degraded_1024.zip` – All 500 normalized degraded images (fixed resolution 1024×128, cropped text‑region only, 500 png files)6. `clean_1024.zip` – All 500 normalized clean images, paired with degraded_1024.zip (fixed resolution 1024×128, cropped text‑region only, 500 png files)7. `txt.zip` – 500 plain‑text transcription files (no side notes, UTF‑8 encoding, line‑level aligned, file names 1.txt to 500.txt correspond to image file names) **Resolution and design rationale:** The full‑resolution images preserve the authentic physical layout of each page:- `degraded/` (original scans): approximately 5000×1000 pixels (varies slightly by page)- `clean/` (reconstructed): approximately 1835×350 pixels (the aspect ratio follows the actual text column height and line length, resulting in small variations in height/width across pages) Both full‑resolution subsets retain the original non‑rectangular page curvature, red bounding frames, and physical edges without any geometric rectification. The clean images are trimmed to the text region plus the red bounding frames, but no cropping of textual content occurs. To facilitate direct use by deep learning models that require fixed‑size inputs, we additionally provide a normalized sub‑set (`degraded_1024/` and `clean_1024/`) at a fixed resolution of 1024×128 pixels. These images are cropped to the main text region (removing marginal page areas such as empty borders or occasional side notes) and anisotropically scaled while preserving the relative geometry of the text layout.This normalized subset is tailored for weak-supervised historical document image restoration, reconstruction, and generation tasks. Users who require the complete page context, including physical edges and original red frames, should use the full-resolution archives instead. All file names are consistent across `degraded1/2`, `clean1/2`, `degraded_1024`, `clean_1024`, and `txt.zip` (1–500). For each image pair (degraded + clean), the file‑naming scheme is identical across all archives (001.png to 500.png). Users downloading degraded1.zip + clean1.zip and degraded2.zip + clean2.zip will obtain the complete set of 500 full‑resolution image pairs. The txt.zip file is provided as a single archive without splitting, as the text files are indexed by file names (1.txt–500.txt) that directly match the image numbers. If you use this dataset in your research, please cite our Data Descriptor paper (currently under review). The associated reconstruction code is available at GitHub: https://github.com/cocotiti123/Dege-Ganjur-Tibetan-Ancient-Book-Helper-Code-and-Datasets/tree/main **Acknowledgements:** We thank the Adarsha (正法宝藏) platform (https://adarshah.org/kangyur/) for providing the original Degé Kangyur images, and the Language Resources Innovation Center of Northwest Minzu University for the curated Tibetan transcriptions. This work was supported in part by the National Natural Science Foundation of China (Grant No. 62466053) ,the Natural Science Foundation of Gansu Province, China (Grant No. 25JRRA993) and the Fundamental Research Funds for the Central Universities(No.31920250082). This dataset is released under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.



