africa-ai-llm-attacks
收藏资源简介:
该数据集是一个合成数据集,专门用于模拟和研究针对非洲国家的AI驱动及大型语言模型(LLM)辅助的网络攻击。随着生成式AI的兴起,非洲的网络威胁格局发生了根本性转变,例如AI生成的钓鱼攻击和深度伪造欺诈在非洲出现了指数级增长。本数据集旨在捕捉这一新兴威胁,包含10,000条平衡记录(50%为攻击,50%为正常),每条记录均标记为合成数据。数据集详细建模了非洲各国的特定攻击模式,例如尼日利亚利用AI生成浪漫骗局脚本和商业邮件欺诈,南非利用AI语音克隆进行银行电话诈骗,肯尼亚利用AI生成招聘骗局内容等。攻击类型覆盖广泛,包括AI生成的钓鱼邮件、LLM辅助的商业邮件欺诈、AI语音克隆、深度伪造视频欺诈、AI生成的浪漫骗局、自动化社会工程学、AI生成的恶意软件等。数据集还模拟了攻击者使用的各种AI工具,包括被滥用的主流LLM(如ChatGPT)、暗网专用网络犯罪LLM(如WormGPT、FraudGPT)、语音克隆API和深度伪造生成工具。攻击内容涵盖了多种非洲本地语言,如英语、法语、斯瓦希里语、豪萨语、约鲁巴语、祖鲁语、阿拉伯语、阿姆哈拉语等。数据集的字段非常丰富,包括记录ID、国家、攻击类型、使用的AI工具、目标行业、语言、传递渠道等基础信息,以及一系列二进制特征,用于标识AI生成的内容模态(文本、语音、图像、视频、代码)、攻击自动化程度、个性化程度、是否使用本地语言、文化适应度、是否绕过传统或AI检测、攻击者技能水平、攻击规模、成功率、造成的财务损失、数据窃取情况、是否被检测到等。此外,还从原始字段中提取了复合特征,如AI模态计数、AI能力评分、语言利用评分、规避评分、攻击威胁评分、犯罪民主化风险评分和检测挑战评分等。该数据集适用于表格分类任务,特别是用于训练和评估模型以检测、分类和分析AI赋能的网络攻击,对于网络安全研究、威胁情报分析和防御策略开发具有重要价值。
This dataset is a synthetic dataset specifically designed to simulate and study AI-driven and large language model (LLM)-assisted cyber attacks targeting African countries. With the rise of generative AI, the cyber threat landscape in Africa has undergone a fundamental shift, such as exponential growth in AI-generated phishing attacks and deepfake fraud in Africa. This dataset aims to capture this emerging threat, containing 10,000 balanced records (50% attacks, 50% normal), each labeled as synthetic data. The dataset models specific attack patterns in various African countries in detail, such as Nigeria using AI to generate romance scam scripts and business email fraud, South Africa using AI voice cloning for bank phone scams, and Kenya using AI to generate recruitment scam content. Attack types cover a wide range, including AI-generated phishing emails, LLM-assisted business email fraud, AI voice cloning, deepfake video fraud, AI-generated romance scams, automated social engineering, AI-generated malware, etc. The dataset also simulates various AI tools used by attackers, including abused mainstream LLMs (e.g., ChatGPT), dark web-specific cybercrime LLMs (e.g., WormGPT, FraudGPT), voice cloning APIs, and deepfake generation tools. Attack content covers multiple African local languages, such as English, French, Swahili, Hausa, Yoruba, Zulu, Arabic, Amharic, etc. The dataset has rich fields, including basic information like record ID, country, attack type, AI tools used, target industry, language, delivery channel, etc., as well as a series of binary features to identify AI-generated content modalities (text, voice, image, video, code), attack automation level, personalization level, whether local language is used, cultural adaptation, whether bypassing traditional or AI detection, attacker skill level, attack scale, success rate, financial loss caused, data theft, whether detected, etc. Additionally, composite features are extracted from original fields, such as AI modality count, AI capability score, language utilization score, evasion score, attack threat score, crime democratization risk score, and detection challenge score. This dataset is suitable for tabular classification tasks, particularly for training and evaluating models to detect, classify, and analyze AI-enabled cyber attacks, and is valuable for cybersecurity research, threat intelligence analysis, and defense strategy development.





