遇见数据集

Benchmark dataset for safety assessment under blast loading

收藏
Mendeley Data2026-04-09 收录
官方服务:

资源简介:

This dataset provides a benchmark comprising 50 curated assessment tasks designed to evaluate the performance of large language model (LLM)-based multi-agent frameworks. Each item represents an independent task with both simple and complex task descriptions, along with corresponding ground-truth references.

本数据集为一款包含50个精心甄选评估任务的基准测试集,旨在评估基于大语言模型(Large Language Model,LLM)的多智能体框架的性能。每项任务均为独立任务,同时配有简易与复杂两种任务描述,以及对应的真值参考。

二维码
社区交流群
二维码
科研交流群
商业服务