遇见数据集

SenseNova-SI-8M

收藏
魔搭社区2026-07-15 更新2026-07-15 收录
官方服务:

资源简介:

**EN** | [中文](README_CN.md) # SenseNova-SI-8M <a href="https://github.com/OpenSenseNova/SenseNova-SI" target="_blank"> <img alt="Code" src="https://img.shields.io/badge/SenseNova_SI-Code-100000?style=flat-square&logo=github&logoColor=white" height="20" /> </a> <a href="https://arxiv.org/abs/2511.13719" target="_blank"> <img alt="arXiv" src="https://img.shields.io/badge/arXiv-SenseNova_SI-red?logo=arxiv" height="20" /> </a> <a href="https://github.com/EvolvingLMMs-Lab/EASI" target="_blank"> <img alt="Code" src="https://img.shields.io/badge/EASI-Code-100000?style=flat-square&logo=github&logoColor=white" height="20" /> </a> <a href="https://easi.lmms-lab.com/leaderboard" target="_blank"> <img alt="Leaderboard" src="https://img.shields.io/badge/%F0%9F%A4%97%20_EASI-Leaderboard-ffc107?color=ffc107&logoColor=white" height="20" /> </a> > **🚀 This is the official full-scale training dataset of the SenseNova-SI series.** > SenseNova-SI-8M contains **~8.16 million** carefully curated training samples spanning **~2.72 million unique images**, organized under a rigorous taxonomy of spatial capabilities. It is the dataset used to train the recommended released model [**SenseNova-SI-1.1-InternVL3-8B**](https://huggingface.co/sensenova/SenseNova-SI-1.1-InternVL3-8B) and serves as the canonical training corpus for spatial intelligence research with SenseNova-SI. ## Overview Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation models to cultivate spatial intelligence within the **SenseNova-SI family**, built upon established multimodal foundations including visual understanding models (i.e., Qwen3-VL and InternVL3) and unified understanding and generation models (i.e., Bagel). We take a principled approach to constructing high-performing and robust spatial intelligence by systematically curating SenseNova-SI-8M: eight million diverse data samples under a rigorous taxonomy of spatial capabilities. SenseNova-SI demonstrates unprecedented performance across a broad range of spatial intelligence benchmarks, while maintaining strong general multimodal understanding. More importantly, we analyze the impact of data scaling, discuss early signs of emergent generalization capabilities enabled by diverse data training, analyze the risk of overfitting and language shortcuts, present a preliminary study on spatial chain-of-thought reasoning, and validate the potential downstream application. SenseNova-SI is an ongoing project, and this report will be updated continuously. All newly trained multimodal foundation models are publicly released to facilitate further research in this direction. *In the future, SenseNova-SI will be integrated with larger-scale in-house models.* ## Release Information **SenseNova-SI-8M** is the official, full-scale training dataset of the SenseNova-SI series. The recommended model [**SenseNova-SI-1.1-InternVL3-8B**](https://huggingface.co/sensenova/SenseNova-SI-1.1-InternVL3-8B) is trained on this dataset and is our **primary recommended release** for spatial intelligence research and downstream use. The previously released [**SenseNova-SI-800K**](https://huggingface.co/datasets/sensenova/SenseNova-SI-800K) subset is provided **only as a reference for scaling-law analysis**; the model trained on it is not the recommended release. <table> <thead> <tr> <th>Model</th> <th>SI Dataset</th> <th>Avg</th> <th>VSI</th> <th>MMSI</th> <th>MindCube-Tiny</th> <th>ViewSpatial</th> <th>SITE</th> <th>BLINK</th> <th>3DSRBench</th> <th>EmbSpatial</th> </tr> </thead> <tbody> <tr> <td>InternVL3-8B</td><td>-</td> <td>45.7</td><td>42.1</td><td>28.0</td><td>41.5</td><td>38.7</td><td>41.1</td><td>53.5</td><td>44.3</td><td><strong>76.3</strong></td> </tr> <tr> <td>VST-7B-SFT</td><td>VST-P-4.1M</td> <td>50.8</td><td>55.5</td><td>32.5</td><td>39.7</td><td>50.5</td><td>39.7</td><td>61.9</td><td>53.1</td><td>73.7</td> </tr> <tr> <td>Cambrian-S-7B</td><td>VSI-590K</td> <td>45.1</td><td>62.9</td><td>27.1</td><td>37.9</td><td>41.3</td><td>36.1</td><td>37.9</td><td>45.0</td><td>72.8</td> </tr> <tr> <td><a href="https://huggingface.co/sensenova/SenseNova-SI-1.1-InternVL3-8B-800K/">*SenseNova-SI-1.1-InternVL3-8B-800K</a></td> <td><a href="https://huggingface.co/datasets/sensenova/SenseNova-SI-800K/">SenseNova-SI-800K</a> (scaling-law reference)</td> <td>-</td><td>60.9</td><td>36.4</td><td>56.9</td><td>52.5</td><td><strong>47.7</strong></td><td>-</td><td>-</td><td>-</td> </tr> <tr> <td><strong><a href="https://huggingface.co/sensenova/SenseNova-SI-1.1-InternVL3-8B/">SenseNova-SI-1.1-InternVL3-8B</a></strong> (recommended)</td> <td><strong>SenseNova-SI-8M</strong> (this dataset)</td> <td><strong>61.5</strong></td> <td><strong>68.8</strong></td> <td><strong>43.3</strong></td> <td><strong>85.7</strong></td> <td><strong>54.7</strong></td> <td><strong>47.7</strong></td> <td><strong>63.9</strong></td> <td><strong>55.5</strong></td> <td><strong>72.0</strong></td> </tr> </tbody> </table> ## What's in this release (vs. SenseNova-SI-800K) | Aspect | SenseNova-SI-800K | **SenseNova-SI-8M (this release)** | |---|---|---| | Role | Scaling-law reference subset | **Official full training set** | | # Training samples | ~832K | **~8.16M** (≈ 800K × 10) | | # Unique images | ~410K | **~2.72M** | | Total image storage | ~370 GB | **~1.1 TB** | | Image archive | 93 × 4 GB independent zips | **53 independent zips (~21 GB each, last ~7 GB)** | | Annotation file | `SenseNova-SI-800K.jsonl` | `SenseNova-SI-8M.jsonl` | | Viewer preview | 1,000-sample parquet | `SenseNova-SI-8M_1000samples.parquet` | | Recommended trained model | (reference only) | **[SenseNova-SI-1.1-InternVL3-8B](https://huggingface.co/sensenova/SenseNova-SI-1.1-InternVL3-8B)** | All zip archives are **independent** (not split volumes) — each can be extracted on its own and together they reconstruct the full `images/` tree. One-click extraction scripts are included for both Linux/macOS and Windows (see [Download & Extract Images](#download--extract-images)). ## Data format The data is stored in **SenseNova-SI-8M.jsonl** using the JSONL (JSON Lines) format, where each line represents an independent data entry. Each entry is a dictionary organized in the following format, containing three main fields: **`id`**, **`conversations`**, and **`image`**. The `id` serves as a unique identifier for each data sample. The `image` field is a list of strings specifying image paths, all given as paths relative to the root data directory. The `conversations` field is a list of dialogue turns, where each turn is a dictionary with two key-value pairs: `from`, indicating the speaker identity (e.g., human or gpt), and `value`, indicating the textual content. Within `value`, the `<image>` placeholder marks where images are inserted, and the number of `<image>` placeholders matches the number of images listed in the `image` field. ```json { "id": 0, "conversations": [ {"from": "human", "value": "<image>\nuser input <image>\nuser input"}, {"from": "gpt", "value": "assistant output"}, {"from": "human", "value": "<image>\nuser input"}, {"from": "gpt", "value": "assistant output"} ], "image": ["path/to/image1.jpg", "path/to/image2.jpg", "path/to/image3.jpg"] } ``` ## Download & Extract Images The image data is packaged into **53 independent zip files** (`images_part_001.zip` through `images_part_053.zip`; ~21 GB each except the last (~7 GB); total ~1.1 TB). Each zip can be extracted on its own — they are **not split volumes**, so you don't need all parts to extract any one of them. Every zip preserves the full `images/` directory structure and extracting them all to the same destination reconstructs the complete image tree. Two one-click extraction scripts are included in the repo root: **Linux / macOS / Git Bash:** ```bash bash extract_all.sh # extract to the script's parent directory bash extract_all.sh /path/to/dir # extract to a specified directory ``` **Windows PowerShell:** ```powershell .\extract_all.ps1 # extract to the script's parent directory .\extract_all.ps1 -Dest D:\data # extract to a specified directory ``` You can also extract manually with any zip tool (e.g. `unzip`, 7-Zip, WinRAR) — each zip is a standard archive. ### Evaluation After training, you can use [EASI](https://github.com/EvolvingLMMs-Lab/EASI) to evaluate your model on mainstream spatial intelligence benchmarks. EASI supports over 20 spatial intelligence models and more than 10 spatial benchmarks, offering Docker for one-click spatial intelligence evaluation. ## 🖊️ Citation ```bib @InProceedings{sensenova-si, title = {Scaling Spatial Intelligence with Multimodal Foundation Models}, author = {Cai, Zhongang and Wang, Ruisi and Gu, Chenyang and Pu, Fanyi and Xu, Junxiang and Wang, Yubo and Yin, Wanqi and Yang, Zhitao and Wei, Chen and Sun, Qingping and Zhou, Tongxi and Li, Jiaqi and Pang, Hui En and Qian, Oscar and Wei, Yukun and Lin, Zhiqian and Shi, Xuanke and Deng, Kewang and Han, Xiaoyang and Chen, Zukai and Fan, Xiangyu and Deng, Hanming and Lu, Lewei and Pan, Liang and Li, Bo and Liu, Ziwei and Wang, Quan and Lin, Dahua and Yang, Lei}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year = {2026} } ```

提供机构:
maas
创建时间:
2026-05-15
二维码
社区交流群
二维码
科研交流群
商业服务