CMAB-The World's First National-Scale Multi-Attribute Building Dataset
收藏资源简介:
Rapidly acquiring three-dimensional (3D) building data, including geometric attributes like rooftop, height and orientations, as well as indicative attributes like function, quality, and age, is essential for accurate urban analysis, simulations, and policy updates. Current building datasets suffer from incomplete coverage of building multi-attributes. This paper presents the first national-scale Multi-Attribute Building dataset (CMAB) with artificial intelligence, covering 3,667 spatial cities, 31 million buildings, and 23.6 billion m² of rooftops with an F1-Score of 89.93% in OCRNet-based extraction, totaling 363 billion m³ of building stock. We trained bootstrap aggregated XGBoost models with city administrative classifications, incorporating morphology, location, and function features. Using multi-source data, including billions of remote sensing images and 60 million street view images (SVIs), we generated rooftop, height, structure, function, style, age, and quality attributes for each building with machine learning and large multimodal models. Accuracy was validated through model benchmarks, existing similar products, and manual SVI validation, mostly above 80%. Our dataset and results are crucial for global SDGs and urban planning.Data records: A building dataset with a total rooftop area of 23.6 billion square meters in 3,667 natural cities in China, including the attribute of building rooftop, height, structure, function, age, style, colour and quality, as well as the code files used to calculate these data. The deep learning models used are OCRNet, XGBoost, fine-tuned CLIP and Yolo-v8.Reference Format:Zhang, Y., Zhao, H. & Long, Y. CMAB: A Multi-Attribute Building Dataset of China. Sci Data 12, 430 (2025). https://doi.org/10.1038/s41597-025-04730-5.
快速获取三维(3D)建筑数据,包括屋顶、高度、朝向等几何属性,以及功能、质量、建造年代等指示性属性,是开展精准城市分析、模拟与政策更新的核心前提。当前主流建筑数据集普遍存在多属性覆盖不全的痛点。本文首次发布基于人工智能构建的国家级多属性建筑数据集(Multi-Attribute Building Dataset,CMAB),覆盖中国3667个自然城市、3100万栋建筑,总屋顶面积达236亿平方米,基于OCRNet的属性提取F1分数达89.93%,建筑总存量体积达3630亿立方米。本研究通过城市行政区划分类训练装袋集成XGBoost模型,融合形态、区位与功能特征;依托多源数据(含数十亿幅遥感影像与6000万张街景图像(SVIs)),结合机器学习与多模态大模型,为每栋建筑生成屋顶、高度、结构、功能、风格、建造年代与质量等属性信息。通过模型基准测试、现有同类产品对比与人工街景图像核验完成精度验证,多数属性的准确率均高于80%。本数据集与研究成果对落实联合国可持续发展目标(SDGs)及开展城市规划具有重要支撑价值。数据记录:本数据集为覆盖中国3667个自然城市的建筑数据集,总屋顶面积达236亿平方米,包含建筑屋顶、高度、结构、功能、建造年代、风格、色彩与质量等属性,附带用于计算上述数据的代码文件。所采用的深度学习模型包括OCRNet、XGBoost、微调版CLIP与Yolov8。引用格式:Zhang, Y., Zhao, H. & Long, Y. CMAB: A Multi-Attribute Building Dataset of China. Sci Data 12, 430 (2025). https://doi.org/10.1038/s41597-025-04730-5.




