遇见数据集

MI4SE

收藏
Figshare2026-03-30 更新2026-04-28 收录
官方服务:

资源简介:

Replication dataset for the paper:"Where Do Vulnerabilities 'Live' in LLMs? A Mechanistic Interpretability Study with Sparse Autoencoders"We use pretrained sparse autoencoders in multiple LLM architectures to extract internal activations as the LLMs process code snippets containing vulnerabilities. We rank the importance of these activations for distinguishing between vulnerabilities (and specific CWE types) and non-vulnerabilities based on pairwise magnitude differences for different code snippets. We then use the top activations to train classifiers to predict the presence of vulnerabilities. This repository contains the code to run these experiments, notebooks to analyze how the abilities of LLMs to understand vulnerabilities changes across sampling conditions, layer splits, and sparsity levels, and scripts to evaluate classification performance for non-SAE models as comparisons.

创建时间:
2026-03-30
二维码
社区交流群
二维码
科研交流群
商业服务