相关数据集
Agent安全评估数据集
Agent Security Bench (ASB) 是一个面向LLM代理安全性的综合性基准框架,涵盖电子商务、自动驾驶、金融等10个场景,涉及10类代理、400余种工具与任务,集成23种攻击/防御方法及7项评估指标。数据集系统评估代理在多场景下的安全脆弱性与防御效能,支撑对提示注入、越权访问等典型攻击的测试。其主要应用于LLM代理的安全性基准评测,助力研发更鲁棒、可靠的智能代理系统。
库帕思2025-12-22 更新150
open-llm-leaderboard/details_chlee10__T3Q-Platypus-Mistral7B
该数据集是在评估模型chlee10/T3Q-Platypus-Mistral7B时自动生成的,包含63个配置,每个配置对应一个评估任务。数据集由1次运行生成,每次运行的结果作为一个特定的分割,分割名称使用运行的时间戳。train分割始终指向最新的结果。此外,results配置存储了所有运行的聚合结果,用于计算和展示在Open LLM Leaderboard上的聚合指标。
Hugging Face2024-03-12 更新90
semeru/code-code-GeneratingAssertsAbstract
--- dataset_info: features: - name: input dtype: string - name: output dtype: string splits: - name: validation num_bytes: 15586336 num_examples: 15809 - name: train nu
Hugging Face2023-05-30 更新60
Research on development of a universal diagnostic system for stator winding faults of induction motor and PMSM based on transfer learning.
Dane przedstawione w zbiorze opisują badania nad wykorzystaniem uczenia transferowego głębokiej sieci konwolucyjnej do opracowania uniwersalnego systemu diagnostyki uszkodzeń uzwojenia stojana silnika
DataCite Commons2024-03-26 更新110
BenchmarkDP - Text extraction from general documents benchmark dataset
This collection contains the benchmark data used for benchmarking text extraction tools. The data conatins : - List of documents - Ground truth data for each document - Additional
DataCite Commons2025-04-01 更新60



