MMAD
收藏资源简介:
MMAD数据集是由南方科技大学和腾讯优图实验室等机构联合创建的首个用于工业异常检测的多模态大语言模型综合基准。该数据集包含39,672个多选题,基于8,366张工业图像,涵盖了38个工业产品类别和244种缺陷类型。数据集的创建过程结合了GPT-4V生成丰富的语义标注,并通过人工审核确保问题和选项的合理性与准确性。MMAD数据集主要应用于工业质量检测领域,旨在评估和提升多模态大语言模型在工业异常检测任务中的性能,解决传统方法在灵活性和详细报告生成方面的不足。
The MMAD dataset is the first comprehensive benchmark for multimodal large language model (LLM)-based industrial anomaly detection, jointly created by institutions including Southern University of Science and Technology and Tencent YouTu Lab. This dataset contains 39,672 multiple-choice questions based on 8,366 industrial images, covering 38 industrial product categories and 244 defect types. The dataset construction process leverages GPT-4V to generate rich semantic annotations, and ensures the rationality and accuracy of questions and options through manual review. The MMAD dataset is primarily applied in the field of industrial quality inspection, with the goal of evaluating and enhancing the performance of multimodal LLMs in industrial anomaly detection tasks, and addressing the limitations of traditional methods in terms of flexibility and detailed report generation.

- 1MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection南方科技大学、腾讯优图实验室、阿尔伯塔大学、上海交通大学 · 2024年



