Benchmark dataset for safety assessment under blast loading
收藏NIAID Data Ecosystem2026-05-10 收录
官方服务:
资源简介:
This dataset provides a benchmark comprising 50 curated assessment tasks designed to evaluate the performance of large language model (LLM)-based multi-agent frameworks. Each item represents an independent task with both simple and complex task descriptions, along with corresponding ground-truth references.
创建时间:
2025-12-25




