遇见数据集

Amharic Phishing Dataset for Ethiopian Fintech Context

收藏
Zenodo2025-05-01 更新2026-05-26 收录
官方服务:

资源简介:

This dataset comprises 2,540 Amharic SMS, each divided into equal sets of actual ('ham', 1270 instances) and phishing/spam ('spam', 1270 instances). It was specifically designed for the Ethiopian scenario with a focus on threats to the country's emerging Fintech sector, such as products like TeleBirr and CBE. The data was collected from anonymized real user messages, social media incidents, and institutional security teams. The data is for training and testing machine learning models on problems such as phishing detection, spam filtering, and adversarial AI research, which are most useful for low-resource languages such as Amharic. The data is provided in plain CSV format with 'label' and 'message' columns. Messages are similar to actual interaction, such as normal scam behavior (spurious notifications, OTP requests, job offers) and real interaction. All personally identifiable information has been anonymized. This dataset was developed under the research of the QuantumShield cybersecurity framework.

提供机构:
Zenodo
创建时间:
2025-05-01
二维码
社区交流群
二维码
科研交流群
商业服务