遇见数据集

Roman Urdu Phishing SMS Dataset for Smishing Detection Research

收藏
Zenodo2026-08-05 更新2026-08-13 收录
官方服务:

资源简介:

This dataset contains labeled Roman Urdu/English SMS messages for phishing (smishing) detection research. It includes: - 1,000 unique messages used in the primary experiments (595 legitimate, 405 phishing) - - Template IDs and Source_Type categories - Both raw and normalized text versions The dataset was created to study the effect of text normalization on machine learning-based smishing detection in Roman Urdu, a low-resource, code-mixed language. A template-aware evaluation protocol was used to prevent template leakage between train and test sets. Related paper: "Does Text Normalization Improve Machine Learning Detection of Roman Urdu Phishing SMS Messages? An Empirical Study with a Template-Aware Evaluation Protocol"

提供机构:
Zenodo
创建时间:
2026-08-05
二维码
社区交流群
二维码
科研交流群
商业服务