AASia: An Intelligent Framework for Generating Asset Administration Shells from Natural Language
收藏资源简介:
This dataset supports the validation of the AASia framework, an LLM-based pipeline for generating Asset Administration Shell (AAS) instances from unstructured natural language descriptions of industrial assets. The dataset is organized into three main parts. AASIA-Manual-AAS-Modeling-Study contains the artifacts from a comparative study of manual AAS modeling. Three industrial test cases — an electric motor (TC1), a centrifugal pump (TC2), and a filling and capping machine (TC3) — were independently modeled by three graduate-level engineering participants using the AASX Package Explorer. The resulting AAS artifacts are provided in .aasx and JSON formats, together with participant evaluation spreadsheets reporting estimated modeling time, perceived effort, and qualitative observations. AASIA-Automated-AAS-Modeling-Study contains the AAS artifacts generated by AASia for the same three test cases, enabling direct structural comparison with the manually created artifacts. For TC3, two executions are included: one using the original input description and one using a revised version with increased explicitness regarding quantitative ranges and units. AASIA-Robustness-Evaluation contains the artifacts from a broader robustness evaluation of the AASia pipeline under its stabilized configuration. A total of 26 executions were performed across 9 distinct industrial asset categories — storage tanks, pressure sensors, homogenizers, spray dryers, air compressors, control valves, belt conveyors, plate pasteurizers, and mixing reactors — using input descriptions ranging from 15 to 150 words in both Spanish and English. Two LLM backends were evaluated: the Groq API with meta-llama/llama-4-scout-17b-16e-instruct and a local LM Studio deployment with meta-llama/llama-3.1-8b-instruct. All 26 executions produced structurally valid AAS Environment JSON documents and AASX packages. Each execution folder contains the input description, the intermediate structured interpretation, the generated AAS JSON, the exported AASX package, and execution logs. The file dataset_index.csv provides a structured summary of all executions, including asset type, input language, LLM configuration, word count, pipeline status, and number of submodels and properties generated. Failed execution attempts are also included for documentation purposes, with observer notes describing the observed issue and resolution. The complete dataset is intended to support reproducibility of the reported analyses, facilitate independent inspection of the generated artifacts, and enable further research on automated AAS generation approaches.
本数据集用于验证AASia框架——一款基于大语言模型(LLM)的流水线,可从工业资产的非结构化自然语言描述生成资产管理壳(Asset Administration Shell, AAS)实例。本数据集分为三个核心部分。 AASIA-Manual-AAS-Modeling-Study 收录了一项手动AAS建模对比研究的相关产物。三名工程硕士研究生参与者依托AASX包资源管理器(AASX Package Explorer),独立完成了三个工业测试用例的建模:电动机(TC1)、离心泵(TC2)以及灌装机与旋盖一体机(TC3)。最终生成的AAS产物以.aasx和JSON格式提供,同时附带参与者评估表格,记录了预估建模时长、感知建模工作量及定性观测结果。 AASIA-Automated-AAS-Modeling-Study 收录了AASia为上述三个测试用例生成的AAS产物,可与手动创建的建模产物直接开展结构对比。针对测试用例TC3,本次研究包含两次执行结果:一次使用原始输入描述,另一次使用优化后的输入描述版本,该版本明确补充了量化范围与单位信息。 AASIA-Robustness-Evaluation 收录了AASia流水线在稳定配置下开展的更广泛鲁棒性评估的相关产物。本次评估共完成26次执行任务,覆盖9类不同的工业资产:储罐、压力传感器、均质机、喷雾干燥机、空气压缩机、控制阀、带式输送机、板式巴氏杀菌机以及搅拌反应器。输入描述的字数介于15至150词之间,同时包含西班牙语与英语两种版本。本次评估共测试两种大语言模型后端:搭载meta-llama/llama-4-scout-17b-16e-instruct的Groq API,以及部署于本地LM Studio的meta-llama/llama-3.1-8b-instruct。所有26次执行均生成了结构合法的AAS环境JSON文档与AASX包。每个执行任务对应的文件夹均包含输入描述、中间结构化解析结果、生成的AAS JSON文件、导出的AASX包以及执行日志。文件dataset_index.csv为所有执行任务提供了结构化汇总信息,涵盖资产类型、输入语言、大语言模型配置、词数、流水线运行状态以及生成的子模型与属性数量。为便于文档记录,本次评估同时收录了失败的执行尝试,并附带观察者记录,详细描述了观测到的问题与解决方案。 本完整数据集旨在支持已发表分析的可复现性,便于独立审查生成的建模产物,并为自动化AAS生成方法的后续研究提供支撑。




