PrivacyXray: Detecting Privacy Breaches in LLMs through Semantic Consistency and Probability Certainty
收藏官方服务:
资源简介:
This artifact contains the code and dataset used in our paper to analyze and classify privacy leakage behaviors in LLMs. It includes scripts for model fine-tuning, hidden state extraction, and classification, as well as training data.
本研究工件包含本论文用于分析与分类大语言模型(Large Language Models)隐私泄露行为的代码与数据集。 其涵盖模型微调、隐状态提取与分类任务相关脚本文件,以及训练数据集。
提供机构:
Zenodo创建时间:
2025-06-07



