Chat-Audio Attacks (CAA)
收藏资源简介:
Chat-Audio Attacks (CAA) 数据集由悉尼科技大学等机构创建,旨在评估大型语言模型在对抗性音频攻击下的鲁棒性。该数据集包含360组音频攻击数据,总计1680个样本,涵盖四种不同类型的音频攻击:内容攻击、情感攻击、显式噪声攻击和隐式噪声攻击。数据集的创建过程包括从公开的多模态数据集中手动收集音频样本,并通过AzureSpeechSDK代理生成攻击音频。CAA数据集主要应用于语音交互领域,旨在探索和提升大型语言模型在对抗性音频环境中的表现和防御机制。
The Chat-Audio Attacks (CAA) dataset was developed by institutions including the University of Technology Sydney and other relevant research bodies, with the core objective of evaluating the robustness of large language models (LLMs) against adversarial audio attacks. This dataset includes 360 sets of audio attack data, totaling 1680 samples, covering four distinct types of adversarial audio attacks: content attacks, emotion attacks, explicit noise attacks, and implicit noise attacks. The construction process of the CAA dataset involves manually collecting audio samples from publicly available multimodal datasets, and generating adversarial audio samples via AzureSpeechSDK agents. The CAA dataset is primarily applied in the field of speech interaction, aiming to explore and enhance the performance and defense mechanisms of large language models in adversarial audio environments.




