b0sungk1m/tamperbench-quantization-qwen3-4b
收藏资源简介:
该数据集用于评估量化压缩是否可以作为隐式篡改操作符影响大型语言模型(LLM)的安全性。实验基于Qwen/Qwen3-4B模型,通过不同的量化方法(INT8和NF4)和篡改攻击(LoRA微调)来测试模型的安全性和实用性。安全性和实用性的评估分别采用了StrongREJECT风格的评分和MMLU-Pro指标。实验结果表明,量化压缩本身不会显著削弱模型的安全性,也不会放大先前的显式篡改攻击的效果。
This dataset evaluates whether quantization compression can act as an implicit tampering operator affecting the safety of large language models (LLMs). The experiment is based on the Qwen/Qwen3-4B model, testing model safety and utility through different quantization methods (INT8 and NF4) and tampering attacks (LoRA fine-tuning). Safety and utility are assessed using StrongREJECT-style scoring and MMLU-Pro metrics, respectively. The results show that quantization compression does not significantly degrade model safety on its own, nor does it amplify the effects of prior explicit tampering attacks.



