遇见数据集

BenchMIRT-model-statistics

收藏
魔搭社区2026-09-02 更新2026-09-06 收录
官方服务:

资源简介:

Permitted Use: The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Disclaimer: This benchmark measures the latent safety and general reasoning scores of LLMs. The data includes prompts and outputs that may contain biased, toxic, or harmful content. The prompts and outputs were generated using existing benchmarks and third party models, which are subject to the license terms of the original sources. Please refer to the Prompt ID, Model ID, and metadata for source information. This dataset contains the per-model statistics for the BenchMIRT project. Original benchmarks used in this analysis: - BBH: [paper](https://arxiv.org/abs/2210.09261), [data](https://github.com/suzgunmirac/BIG-Bench-Hard) - GPQA: [paper](https://arxiv.org/abs/2311.12022), [data](https://huggingface.co/datasets/Idavidrein/gpqa) - MMLU-Pro: [paper](https://arxiv.org/abs/2406.01574), [data](https://github.com/TIGER-AI-Lab/MMLU-Pro) - MATH: [paper](https://arxiv.org/abs/2103.03874), [data](https://github.com/hendrycks/math/) - MuSR: [paper](https://arxiv.org/abs/2310.16049), [data](https://github.com/Zayne-Sprague/MuSR) - IFEval: [paper](https://arxiv.org/abs/2311.07911), [data](https://github.com/google-research/google-research/tree/master/instruction_following_eval) - BBQ: [paper](https://arxiv.org/abs/2110.08193), [data](https://github.com/nyu-mll/BBQ) - Do-Anything-Now: [paper](https://arxiv.org/abs/2308.03825), [data](https://github.com/verazuo/jailbreak_llms) - HarmBench: [paper](https://arxiv.org/abs/2402.04249), [data](https://github.com/centerforaisafety/HarmBench) - StrongReject: [paper](https://arxiv.org/abs/2402.10260), [data](https://strong-reject.readthedocs.io/en/latest/) - ToxiGen: [paper](https://arxiv.org/abs/2203.09509), [data](https://github.com/microsoft/ToxiGen) - TrustLLM-JailbreakTrigger: [paper](https://arxiv.org/abs/2401.05561v2), [data](https://github.com/HowieHwong/TrustLLM) - WildGuardTest: [paper](https://arxiv.org/abs/2406.18495), [data](https://huggingface.co/datasets/allenai/wildguardmix) - WildJailbreak: [paper](https://arxiv.org/abs/2406.18510), [data](https://huggingface.co/datasets/allenai/wildjailbreak) - WMDP: [paper](https://arxiv.org/abs/2403.03218), [data](https://www.wmdp.ai/) - XSTest: [paper](https://arxiv.org/abs/2308.01263), [data](https://github.com/paul-rottger/xstest)

提供机构:
maas
创建时间:
2026-09-02
二维码
社区交流群
二维码
科研交流群
商业服务