FDM-Bench
收藏资源简介:
FDM-Bench是一个用于评估大型语言模型在熔融沉积建模(FDM)任务中的基准数据集。该数据集由伊利诺伊大学厄巴纳-香槟分校和罗格斯大学等机构创建,旨在通过用户查询和G代码样本评估模型在FDM任务中的表现。数据集包含不同经验水平的用户查询和带有多种异常的G代码样本,帮助评估模型在检测打印缺陷和响应用户查询方面的能力。FDM-Bench的应用领域主要集中在增材制造中的缺陷检测和优化打印质量,旨在解决FDM技术中的复杂参数管理和缺陷诊断问题。
FDM-Bench is a benchmark dataset designed for evaluating the performance of Large Language Models (LLMs) on Fused Deposition Modeling (FDM) tasks. Developed by the University of Illinois Urbana-Champaign, Rutgers University, and other institutions, this dataset aims to assess LLMs' performance on FDM tasks through user queries and G-code samples. It includes user queries from users with varying levels of expertise, as well as G-code samples containing multiple types of anomalies, enabling the evaluation of models' abilities to detect printing defects and respond to user inquiries. The primary application areas of FDM-Bench lie in defect detection and print quality optimization within additive manufacturing, and it is intended to address the challenges of complex parameter management and defect diagnosis in FDM technology.
FDM-Bench 数据集概述
FDM-Bench 是一个用于评估大型语言模型(LLMs)在熔融沉积建模(FDM)特定任务上的基准数据集,包括 G-code 异常检测、用户查询和多项选择题。
数据集概览
1. 带标签的 G-Codes
- 异常检测:包含标记了特定异常的 G-codes,包括无缺陷(Non-defective, ND)、欠挤出(Under-extrusion, UE)、过挤出(Over-extrusion, OE)和面条状缺陷(Spaghetti, SP)。
- 评估类型:每个 G-code 可用于确定性标签(单一标签输出)和基于概率的标签。
2. 自由形式问题
- 专业水平:问题按用户专业水平分类,包括初学者(Beginner)、有经验者(Experienced)和理论水平(Theoretical)。
- 示例:提供参考以确保评估一致性。
- 评估者评分:提供四个 LLM 模型(GPT-4、Claude、Llama-3.1-70B 和 Llama-3.1-405B)在三个指标(准确性、精确性和相关性)上的评分,评分范围为 1 到 5。
3. 多项选择题(MCQs)
- 经验水平:与自由形式问题类似,分为初学者(B)、有经验者(E)和理论水平(T)。
- 标准答案:包含答案键以评估模型准确性。
4. 提示
- 任务特定提示:包括每种任务类型的提示(G-code 检测、自由形式响应和多项选择题)。
如需更详细的信息,请参阅相关研究论文。




