WeldLLM: A Multimodal Framework for Welding Defect Detection and Automated Diagnostic Reasoning
收藏资源简介:
Due to the diverse morphology, low contrast, and complex spatial distribution of weld defects, traditional vision-based detection methods exhibit significant limitations in accuracy and robustness. To address these challenges, this study proposes WeldLLM, a multimodal vision-language framework integrating precise defect localization with domain-specific diagnostic reasoning. First, we develop a customized detection network, WTDWR-YOLO, tailored specifically for welding scenarios. Within this network, we introduce a Dilated Weighted Residual Segmentation module to capture multi-scale spatial features of defects, and construct a Wavelet-Transformed Convolution module to enhance semantic representation in the frequency domain, thereby significantly improving the network’s ability to identify low-contrast defects. Subsequently, we design a lightweight modality alignment module, which encodes regional visual features and structured detection attributes into a unified token sequence. This sequence is then processed by the pre-trained vision-language model Qwen-VL for high-level semantic reasoning. To efficiently adapt the model to the welding domain, we employ Low-Rank Adaptation (LoRA) for rapid fine-tuning, enabling expert-level diagnostic reasoning and the generation of interpretable visual reports. Experimental results on real-world weld defect datasets demonstrate that WeldLLM significantly outperforms existing methods in terms of detection accuracy and interpretability. This research validates the effectiveness and practicality of multimodal vision-language models for automated welding quality assessment, underscoring their high potential for engineering applications.



