CryptoBench: A benchmark of cryptographic API misuse for evaluating LLM code reviewers
收藏官方服务:
资源简介:
54 Python snippets covering nine cryptographic misuse classes (36 vulnerable, 18 matched secure controls), a self-contained harness that queries local Ollama models with a fixed verdict-only prompt, and 1,890 recorded trials across seven open code models (qwen2.5-coder 0.5B to 14B, deepseek-coder 6.7B, codellama 7B). Includes an analysis script that reproduces per-model and per-class detection rates with Wilson 95% confidence intervals.
提供机构:
Zenodo创建时间:
2026-09-30



