IFC, Uniclass
收藏资源简介:
本研究使用的数据集包括IFC和Uniclass,这两个数据集分别由buildingSMART International和国家建筑规范(NBS)维护。IFC数据集提供了建筑和基础设施项目的全面数字描述,而Uniclass则是一个统一的建筑环境分类系统,涵盖了超过8000种产品类型。数据集的创建过程包括从原始数据源中提取产品名称、描述和标签,并通过生成语言模型进行数据增强和校对。这些数据集主要用于评估预训练文本嵌入模型在建筑资产信息管理中的对齐效果,旨在解决建筑资产数据的多源性和多学科性带来的对齐挑战。
The datasets utilized in this study include IFC and Uniclass, which are maintained by buildingSMART International and the National Building Specification (NBS) respectively. The IFC dataset provides comprehensive digital descriptions of construction and infrastructure projects, while Uniclass is a unified built environment classification system covering over 8,000 product types. The development of these datasets involves extracting product names, descriptions and tags from their original data sources, followed by data augmentation and proofreading using generative language models. These datasets are primarily employed to evaluate the alignment performance of pre-trained text embedding models in built asset information management, aiming to address the alignment challenges brought by the multi-source and multi-disciplinary nature of built asset data.

- 1Benchmarking pre-trained text embedding models in aligning built asset information高等技术学院 · 2024年



