遇见数据集

UzSpam v1.1: An Uzbek-language SMS and e-mail spam corpus

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

1. Motivation Why was the dataset created? Uzbek is a less-resourced language with no publicly available labelled corpus for spam and phishing detection. UzSpam was created to support the development and evaluation of spam filters for Uzbek-language SMS, e-mail and messenger messages. It was developed as part of a doctoral research project. Who created it and who funded it? The corpus was created by the author. Funding: to be completed (or "no dedicated funding"). 2. Composition What do instances represent? Each instance is one short text message (SMS, e-mail body or messenger post) with a binary label: 0 = legitimate (ham), 1 = spam (advertising, fraud or phishing).

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务