difraud/difraud
收藏资源简介:
DIFrauD(领域独立欺诈检测基准)是一个包含95,854个样本的标注语料库,涵盖了七个独立领域的欺骗和真实文本。这些领域包括钓鱼邮件、假新闻、政治声明、产品评论、招聘诈骗、短信和Twitter谣言。每个任务都被转换为一个二元分类问题,其中y是欺骗的指示器。数据集经过清理和标准化处理,确保文本的有效性和一致性。
DIFrauD (Domain-Independent Fraud Detection Benchmark) is an annotated corpus consisting of 95,854 samples, covering deceptive and genuine texts from seven independent domains. These domains include phishing emails, fake news, political statements, product reviews, recruitment scams, short message service (SMS) texts, and Twitter rumors. Each task is formulated as a binary classification task, where y serves as the indicator of deception. The dataset has been cleaned and standardized to ensure the validity and consistency of the texts.
DIFrauD - Domain Independent Fraud Detection Benchmark
数据集概述
基本信息
- 语言: 英语
- 许可证: MIT
- 数据量: 10K < n < 100K
- 多语言性: 单语种
- 任务类别: 文本分类, 零样本分类
- 标签: 欺诈检测, 欺骗检测, 钓鱼, 假新闻, 基准, 意见垃圾邮件, 多领域
数据集详情
- 总样本数: 95,854
- 欺骗性样本数: 37,282
- 非欺骗性样本数: 58,572
- 领域数量: 7
领域分类
-
Phishing
- 总数: 15,272
- 欺骗性: 6,074
- 非欺骗性: 9,198
-
Fake News
- 总数: 20,456
- 欺骗性: 8,832
- 非欺骗性: 11,624
-
Political Statements
- 总数: 12,497
- 欺骗性: 8,042
- 非欺骗性: 4,455
-
Product Reviews
- 总数: 20,971
- 欺骗性: 10,492
- 非欺骗性: 10,479
-
Job Scams
- 总数: 14,295
- 欺骗性: 599
- 非欺骗性: 13,696
-
SMS
- 总数: 6,574
- 欺骗性: 1,274
- 非欺骗性: 5,300
-
Twitter Rumours
- 总数: 5,789
- 欺骗性: 1,969
- 非欺骗性: 3,820
数据文件结构
每个领域/数据集分为三个文件:
train.jsonltest.jsonlvalidation.jsonl
数据分割
- 训练集: 80%
- 测试集: 10%
- 验证集: 10%
- 采样过程: 随机,种子=42,按标签分层
数据字段
每个jsonl文件包含两个字段:
text(字符串)label(整数)
label字段表示文本是否欺骗性:
1: 欺骗性0: 非欺骗性
数据处理和清洗
- 使用Cleanlab进行清洗
- 移除非英语条目、解析错误条目、空条目、重复条目
- 条目长度限制在2到1,000,000字符之间
- 标签手动校正
- 空白、引号、项目符号、Unicode标准化
许可证
- MIT许可证




