DL3DV-2k
收藏资源简介:
# DL3DV-2K 📖<a href="https://arxiv.org/abs/2605.23897">Paper</a> | 🏠<a href="https://github.com/InternLM/ETCHR">Homepage</a > | 🤗<a href="https://huggingface.co/internlm/ETCHR-FLUX.2-klein-9B">ETCHR-FLUX.2-klein-9B Model</a > | 🤗<a href="https://huggingface.co/datasets/BeichenZhang/ETCHR-SFT-400K">ETCHR SFT-400K Dataset</a > | 🤗<a href="https://huggingface.co/datasets/internlm/ETCHR-GRPO-10K">ETCHR GRPO-10K Dataset</a > | 🤗<a href="https://huggingface.co/datasets/internlm/DL3DV-2k">DL3DV-2K Benchmark</a > DL3DV-2K is a benchmark constructed from the DL3DV dataset for evaluating the viewpoint transformation capability of large models in spatial reasoning tasks, comprising 2K samples in total. Each sample contains: images (original images), aux_images (transformed images provided for human reference only and not used as question input), question, candidates, and answer. The model needs to imagine the viewpoint of the aux_images from the images in order to effectively answer the question. ## ✒️Citation If you find this project useful, please kindly cite: ``` @article{zhang2026etchr, title={ETCHR: Editing To Clarify and Harness Reasoning}, author={Beichen Zhang, Yuhong Liu, Jinsong Li, Yuhang Zang, Jiaqi Wang, Dahua Lin}, journal={arXiv preprint arXiv:2605.23897}, year={2026} } ```



