P4Ms-sqa
收藏资源简介:
P4Ms基准数据集是用于分析大型多模态模型(LMMs)在多阶段和多模态训练管道中隐私泄露问题的第一个基准。这个特定的数据集代表了训练管道的一个阶段,即语音问答(SQA)。它包含音频记录与问题和答案的配对,模型需要从音频中提取敏感信息。数据集由1,500个独特个体的合成数据组成,涵盖文本、视觉和语音模态。数据集分为三个子集:members(包含用于目标模型训练的个人身份信息PII的样本)、nonmembers(包含在训练期间保留的PII样本)和wo_piis(Without PIIs,不包含任何PII但结构和上下文与敏感样本相同的样本)。数据集的任务类型为SQA,敏感信息包括共享的PII(如姓名、电子邮件、电话号码、信用卡号)和阶段特定的PII(如出生日期和出生地)。音频是通过Chatterbox TTS合成的,使用了Common Voice数据集中的声音。
The P4Ms benchmark dataset is the first benchmark designed to analyze privacy leakage issues of large multimodal models (LMMs) in multi-stage and multimodal training pipelines. This specific dataset represents one stage of the training pipeline, namely Speech Question Answering (SQA). It consists of pairs of audio recordings paired with questions and answers, where models are required to extract sensitive information from the audio. The dataset is composed of synthetic data from 1,500 unique individuals, covering text, visual, and speech modalities. The dataset is divided into three subsets: members (samples containing personally identifiable information (PII) used for target model training), nonmembers (samples containing PII that was withheld during training), and wo_piis ("Without PIIs": samples that do not contain any PII but have the same structure and context as sensitive samples). The task type of the dataset is SQA, and sensitive information includes shared PII such as names, email addresses, phone numbers, and credit card numbers, as well as stage-specific PII such as date of birth and place of birth. The audio was synthesized via Chatterbox TTS using voices from the Common Voice dataset.
P4Ms-sqa 数据集概述
数据集来源
- 数据集名称: P4Ms-sqa
- 所属基准: P4Ms Benchmark (Privacy Measurements for Multistage and Multimodal Models)
- 关联论文: "P4Ms: Privacy Measurements for Multistage and Multimodal Models"
数据集描述
- 核心目的: 作为P4Ms基准的一部分,旨在分析大型多模态模型(LMMs)在现实、多阶段、多模态训练管道中的隐私泄露问题。
- 本数据集定位: 代表训练管道中的一个特定阶段——语音问答(SQA)。
- 内容: 包含音频录音与问答对的配对数据,模型需要从音频中提取敏感信息。
- 数据生成:
- 使用GPT-4.1生成文本转录。
- 使用Chatterbox TTS(https://www.resemble.ai/chatterbox/)进行语音合成。
- 合成语音使用来自Common Voice数据集的音色。
数据集结构
数据特征
user_id: (字符串)path: (音频)conversation: (列表)instruction: (字符串)output: (字符串)
数据子集划分
数据集划分为三个独立的子集:
members:- 描述: 包含在目标模型训练期间使用的、含有个人可识别信息(PII)的样本。
- 样本数量: 4438
- 数据大小: 6696229387.69 字节
non_members:- 描述: 包含在训练期间被保留的、含有PII的样本。
- 样本数量: 5963
- 数据大小: 5002234276.494 字节
without_piis:- 描述: 遵循与敏感样本相同的结构和上下文,但不包含任何PII的样本。用于训练模型学习上下文而不暴露敏感数据。
- 样本数量: 1896
- 数据大小: 2544649859.728 字节
整体统计
- 下载大小: 13908385103 字节
- 数据集总大小: 14243113523.912 字节
数据集摘要
- 阶段: 多模态适应 - 语音
- 类型: 语音 + 文本(问答对)
- 任务: 语音问答(SQA)
- 简短描述: 语音录音,内容为个人自我介绍并提及个人详细信息。
- 敏感信息:
- 共享PII: 姓名、电子邮件、电话号码、信用卡号。
- 阶段特定PII: 出生日期和出生地。
技术信息
- 任务类别: 问答
- 语言: 英语
- 标签: 音频、语音




