遇见数据集

ImageNet-1K-Camera

收藏
魔搭社区2026-07-19 更新2026-07-19 收录
官方服务:

资源简介:

# ImageNet-1K-Camera ![camera map collage](analysis/ImageNet-1K-Camera.png) Per-image camera parameter annotations for the full **ImageNet-1K** dataset (1,000 training classes + the 50,000-image validation split, ~1.35M images), captioned by the [**Puffin-World**](https://github.com/KangLiao929/Puffin) model. More captioned datasets are provided in our [**Puffin-16M**](https://kangliao929.github.io/projects/puffin-16m/) website. The collage above visualizes the camera maps on sample images — each pair shows the **up field** (green arrows: the projected gravity-up direction) and the **latitude field** (colored contours: angle above/below the horizon). ## Format The archive mirrors the source ImageNet WebDataset layout: one `.tar` per source shard (`n*.tar` per class, plus `val.tar`), each containing one `.json` per image whose name matches the source image stem. Each JSON holds the predicted monocular camera parameters: | Field | Meaning | Unit | |-------|---------|------| | `roll` | camera roll | radians | | `pitch` | camera pitch | radians | | `vfov` | vertical field-of-view | radians | | `k1` | radial distortion coefficient | – | | `parse_ok` | whether the model output parsed within valid ranges | bool | Example: ```json {"roll": 0.0123, "pitch": -0.0871, "vfov": 1.0123, "k1": 0.0000, "parse_ok": true} ``` ## Camera Parameter Distributions Histograms of the predicted roll / pitch / vertical-FoV, **train and val plotted separately** (proportion of valid samples per 10° bin; `parse_ok=False` excluded). ### Train (1,278,950 images) ![train camera stats](analysis/imagenet1k_train_camera_stats.png) ### Val (49,924 images) ![val camera stats](analysis/imagenet1k_val_camera_stats.png) | split | roll μ / med / σ | pitch μ / med / σ | FoV μ / med / σ | |-------|------------------|-------------------|-----------------| | train | 0.1° / 0.0° / 7.8° | −7.8° / −3.7° / 15.7° | 28.9° / 25.6° / 8.9° | | val | 0.1° / 0.0° / 8.5° | −8.8° / −4.5° / 16.7° | 30.3° / 27.5° / 9.1° | - **Roll** is sharply peaked at 0° (images shot upright/level). - **Pitch** is slightly negative (a mild downward-looking tendency). - **FoV** concentrates in 20–40° (median ≈ 26–28°); train and val agree closely. If you'd like a dataset with a more diverse and uniform distribution of camera parameters, please refer to our [Puffin-4M](https://huggingface.co/datasets/KangLiao/Puffin-4M) and [Puffin-16M](https://huggingface.co/datasets/KangLiao/Puffin-16M) datasets. ### Dataset Download You can download the entire dataset using the following command: ```bash hf download KangLiao/ImageNet-1K-Camera --repo-type dataset ``` ### Caption Pipeline Beyond this captioned dataset, we also release **a complete captioning pipeline** for annotating camera parameters for arbitrary datasets, analyzing camera parameter distributions, and visualizing the corresponding camera maps. The pipeline is available in our [GitHub repository](https://github.com/KangLiao929/Puffin). ### Citation If you find the captioned dataset useful for your research or applications, please cite our paper using the following BibTeX: ```bibtex @article{liao2025puffin, title={Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation}, author={Liao, Kang and Wu, Size and Wu, Zhonghua and Jin, Linyi and Wang, Chao and Wang, Yikai and Wang, Fei and Li, Wei and Loy, Chen Change}, journal={arXiv preprint arXiv:2510.08673}, year={2025} } ```

提供机构:
maas
创建时间:
2026-06-25
二维码
社区交流群
二维码
科研交流群
商业服务