遇见数据集

kth8/gpt-oss-20b-MMLU-Pro-benchmark

收藏
Hugging Face2026-04-10 更新2026-04-12 收录
官方服务:

资源简介:

--- license: apache-2.0 language: - en base_model: openai/gpt-oss-20b datasets: - TIGER-Lab/MMLU-Pro --- Benchmark of [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b) against [TIGER-Lab/MMLU-Pro](https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro) dataset. Accuracy: 75.5% with Python tool. | Metric | Value | |----------------------|---------------| | **Correct** | 755 | | **Incorrect** | 245 | | **Errors** | 0 | | **Total samples** | 1000 | | **Python tool calls**| 411 | | **Total completion tokens** | 1,179,150 | Raw stats: ```json { "accuracy": 0.755, "correct": 755, "incorrect": 245, "error": 0, "total": 1000, "python_tool_calls": 411, "completion_tokens": 1179150 } ```

许可证:Apache-2.0 语言: - en(英语) 基础模型:openai/gpt-oss-20b 测试数据集: - TIGER-Lab/MMLU-Pro 本基准测试针对[openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b) 在[TIGER-Lab/MMLU-Pro](https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro)数据集上的性能表现展开。使用Python工具测得的准确率为75.5%。 | 评估指标 | 数值 | |----------------------|---------------| | **正确样本数** | 755 | | **错误样本数** | 245 | | **异常错误数** | 0 | | **总样本数** | 1000 | | **Python工具调用次数**| 411 | | **总补全Token(Completion Tokens)数** | 1,179,150 | 原始统计数据: json { "accuracy": 0.755, "correct": 755, "incorrect": 245, "error": 0, "total": 1000, "python_tool_calls": 411, "completion_tokens": 1179150 }

提供机构:
kth8
二维码
社区交流群
二维码
科研交流群
商业服务