GovSecLLM++ SECAI 2026 Artifact Package
收藏资源简介:
This artifact package accompanies the SECAI 2026 submission: **GovSecLLM++: A Compliance-Aware Benchmark for Security Testing and Governance Evidence in LLM-Based Applications** GovSecLLM++ is a compliance-aware benchmark and evaluation protocol for security testing of LLM-based applications. The benchmark maps governance requirements to application-level risks, adversarial scenario cards, expected safe behaviours, measurable security and governance metrics, and audit evidence. This package contains the experimental artifacts without the final paper source/PDF. It includes: - AdaptiveGovSec-120 outputs and summaries; - corrected F2 hidden-context confidentiality rerun outputs; - mini multi-model benchmark outputs; - strict rescoring files; - human-validation summary workbook; - scoring utilities; - reproducibility notes; - data card; - checksums and package manifest. The main experimental components are: 1. Controlled validation with 6300 runs. 2. Hosted Groq full-suite evaluation with 1050 real-backend runs. 3. Multi-model mini-benchmark and strict rescoring. 4. Corrected F2 rerun to separate hidden-context leakage from user-provided canary echoing. 5. AdaptiveGovSec-120 with 478 successful calls out of 480 expected calls. 6. Human-confirmed validation on a stratified sample of 50 outputs. All secrets, policy records, documents, tools, and credentials are synthetic. The artifact does not contain real personal data, real credentials, or real confidential organizational data. The benchmark provides technical and governance-oriented evaluation evidence. It does not constitute legal certification of compliance with the EU AI Act, ISO/IEC 42001, or any other regulatory framework.



