官方服务:
资源简介:
Data and source code for Loegised and the baseline methods.
应用场景:
创建时间:
2025-04-08
相关数据集
CRMArenaPro
# Dataset Card for CRMArena-Pro - [Dataset Description](https://huggingface.co/datasets/Salesforce/CRMArenaPro/blob/main/README.md#dataset-description) - [Paper Information](https://huggingface.co/
魔搭社区2026-04-28 更新220
OMNIX_Benchmarks_Latest
OMNIX Benchmarks 是一个用于评估和基准测试大型语言模型性能的数据集。它旨在通过一套标准化的指标对模型进行综合比较,涵盖格式遵循、逻辑推理、知识回忆、约束遵循、成功率(首次通过和最终)以及推理延迟等多个关键维度。该数据集支持文本生成和问答任务类别,并提供了对多个流行模型(如 Qwen、Gemma、Llama 等系列的不同版本)在上述维度上的量化评分和排名分析,帮助研究人员和开发者了解
Hugging Face2026-07-07 更新100
open-llm-leaderboard/details_chlee10__T3Q-Platypus-Mistral7B
该数据集是在评估模型chlee10/T3Q-Platypus-Mistral7B时自动生成的,包含63个配置,每个配置对应一个评估任务。数据集由1次运行生成,每次运行的结果作为一个特定的分割,分割名称使用运行的时间戳。train分割始终指向最新的结果。此外,results配置存储了所有运行的聚合结果,用于计算和展示在Open LLM Leaderboard上的聚合指标。
Hugging Face2024-03-12 更新90
DCAgent2/dev_set_v2_100k_epochs3__Qwen3_8B_20260325_231128
--- dataset_info: features: - name: conversations list: - name: content dtype: string - name: role dtype: string - name: agent dtype: string - name: model dtype
Hugging Face2026-03-26 更新110
smalleval/mmlu-nano
# SmallEval: Browser-Friendly LLM Evaluation Datasets 🚀 [](https://cloudcode.ai) SmallEval is a curated
Hugging Face2025-01-20 更新90



