Post-Quantum-Cryptography-Benchmark
收藏资源简介:
一个专门用于评估LLMs在后量子密码学方面能力的基准测试数据集,包括1695个单选题、565个多选题和428个开放式问答对。该数据集涵盖了后量子密码学的八个主要分支:数学基础、安全归约理论、核心算法设计与构造、标准化与密钥封装机制、算法软件与硬件实现与优化、密码协议迁移、侧信道攻击与保护、PQC迁移与混合部署策略。
A benchmark dataset specifically developed for evaluating the capabilities of Large Language Models (LLMs) in the field of post-quantum cryptography (PQC). It comprises 1,695 single-choice questions, 565 multiple-choice questions, and 428 open-ended question-answer pairs. This dataset covers eight core branches of post-quantum cryptography: foundational mathematics, security reduction theory, core algorithm design and construction, standardization and key encapsulation mechanisms (KEMs), software and hardware implementation and optimization of cryptographic algorithms, cryptographic protocol migration, side-channel attacks and countermeasures, and PQC migration and hybrid deployment strategies.
数据集概述
数据集名称
Post-Quantum-Cryptography-Benchmark
数据集简介
该数据集是一个专门用于评估大型语言模型在后量子密码学领域能力的基准测试集。它包含三种类型的问题对:
- 1695个单项选择题对。
- 565个多项选择题对。
- 428个开放式问答对。
覆盖领域
数据集涵盖了后量子密码学的八个主要分支:
- 后量子密码学的数学基础
- 安全归约理论
- 核心算法设计与构造
- 标准化与密钥封装机制
- 算法的软件与硬件实现及优化
- 密码协议迁移
- 侧信道攻击与防护
- 后量子密码迁移与混合部署策略
模型测试结果
数据集创建者对17个模型(包括16个主流大型语言模型和一个名为PQC-LLM的微调模型)进行了测试,性能排名如下:
| 排名 | 模型 | 问答(基础) | 问答(扩展) | 多项选择(F1) | 多项选择(EM) | 单项选择(Acc) |
|---|---|---|---|---|---|---|
| 1 | PQC-LLM | 74.48 | 76.1 | 89.4 | 52.9 | 86.1 |
| 2 | Qwen3-235B | 75.5 | 71.9 | 91.1 | 48.1 | 80.1 |
| 3 | Claude-opus | 72.4 | 68.8 | 88.3 | 46.2 | 83.2 |
| 4 | Gemini-3-flash | 71.4 | 73.9 | 87.8 | 43.4 | 80.7 |
| 5 | GLM-4.7 | 70.5 | 72.8 | 87.0 | 41.6 | 81.8 |
| 6 | Mistral-large | 70.3 | 67.2 | 88.1 | 45.8 | 77.6 |
| 7 | Doubao | 74.5 | 68.7 | 86.7 | 42.5 | 75.3 |
| 8 | GPT-5.2 | 74.1 | 70.7 | 84.6 | 38.6 | 77.2 |
| 9 | DeepSeek-V3.2 | 73.7 | 69.7 | 84.4 | 38.1 | 78.1 |
| 10 | LongCat | 70.0 | 64.4 | 87.3 | 42.7 | 78.9 |
| 11 | HY | 70.8 | 66.8 | 86.2 | 38.6 | 78.6 |
| 12 | ERNIE | 70.1 | 66.6 | 83.9 | 40.7 | 73.8 |
| 13 | MiniMax | 66.8 | 70.2 | 83.6 | 37.2 | 70.7 |
| 14 | llama-4-maverick | 65.9 | 67.0 | 85.1 | 36.6 | 76.3 |
| 15 | Ling | 67.1 | 67.5 | 81.6 | 32.7 | 74.0 |
| 16 | MiMo-v2 | 68.3 | 65.7 | 83.6 | 37.2 | 61.4 |
| 17 | Grok-4-1 | 65.2 | 63.5 | 83.0 | 29.6 | 64.4 |
数据集地址
https://github.com/zhe-liangzhi/Post-Quantum-Cryptography-Benchmark




