LabSafety Bench
收藏资源简介:
LabSafety Bench是由圣母大学和IBM研究团队创建的一个实验室安全评估框架,旨在评估大型语言模型(LLMs)在实验室安全环境中的可靠性。数据集包含765个多选题,涵盖了实验室安全的四个主要领域:危险物质、应急响应、责任与合规、设备与材料处理。这些问题由人类专家验证,确保其准确性和清晰性。数据集的创建过程包括提出新的实验室安全分类法、收集相关材料、生成问题并由GPT-4o协助优化选项,最后由专家审核。该数据集主要用于评估LLMs在实验室安全决策中的应用,旨在提高实验室环境中的安全性和可靠性。
LabSafety Bench is a laboratory safety evaluation framework developed by teams from the University of Notre Dame and IBM Research, which aims to assess the reliability of Large Language Models (LLMs) in laboratory safety contexts. The dataset contains 765 multiple-choice questions spanning four core domains of laboratory safety: hazardous substances, emergency response, responsibility and compliance, and equipment and material handling. All questions have been verified by human experts to ensure their accuracy and clarity. The dataset creation process includes proposing a novel laboratory safety taxonomy, collecting relevant materials, generating questions, optimizing the options with the assistance of GPT-4o, and finally conducting expert reviews. This dataset is primarily used to evaluate the application of LLMs in laboratory safety decision-making, with the goal of enhancing safety and reliability in laboratory environments.




