CaptureGuide-Dataset
收藏资源简介:
CaptureGuide-Dataset是由复旦大学与StepFun联合构建的大规模摄影指导数据集,旨在支持实时拍摄过程中的构图决策与姿态推荐研究。该数据集共包含约13万样本,涵盖摄影师侧的10万构图指导样本和主体侧的3万姿态指导样本,数据源自在线平台图像,并经过专家标注与多模态大语言模型验证的自我蒸馏流程处理。数据集通过结构化标注流程构建,包括专家标注的种子集、MLLM验证的伪标签扩展以及人机协同的理性生成,确保了数据的高质量与一致性。该数据集主要应用于摄影美学、多模态人工智能及人机交互领域,旨在解决现有裁剪基准忽略实时拍摄引导与主体姿态建议的局限,为开发能够提供交互式拍摄辅助的智能模型提供数据支撑。
CaptureGuide-Dataset is a large-scale photography guidance dataset jointly constructed by Fudan University and StepFun, aiming to support research on composition decision-making and pose recommendation during real-time shooting. This dataset contains approximately 130,000 samples in total, including 100,000 composition guidance samples from the photographer's perspective and 30,000 pose guidance samples from the subject's perspective. The data is sourced from images on online platforms, and processed via expert annotation and a self-distillation workflow validated by multi-modal large language models (MLLMs). The dataset is built through a structured annotation pipeline, which consists of an expert-annotated seed set, pseudo-label expansion verified by MLLMs, and human-machine collaborative rational generation, ensuring high data quality and consistency. This dataset is mainly applied in the fields of photographic aesthetics, multimodal artificial intelligence, and human-computer interaction. It aims to address the limitation that existing cropping benchmarks ignore real-time shooting guidance and subject pose suggestions, providing data support for developing intelligent models capable of delivering interactive shooting assistance.





