TeleAntiFraud 2.0
收藏资源简介:
TeleAntiFraud 2.0是一个可刷新的音频基准数据集,专为电信诈骗检测设计,由中国人民公安大学等机构联合创建。当前版本包含两个月度快照,共1800条中文电话通话,其中1200条为欺诈样本,600条为近域非欺诈样本。数据集采用混合树反诈骗生成流水线,将在线诈骗摘要转化为结构化的角色匹配语音对话,确保欺诈与合法路径共享上下文仅在关键行为处分叉。该基准旨在解决现有基准无法持续更新欺诈模式且缺乏近域负样本的问题,为评估模型在真实混淆条件下的音频诈骗检测性能提供更严格的测试平台。
TeleAntiFraud 2.0 is a refreshable audio benchmark dataset purpose-built for telecom fraud detection, jointly developed by the People's Public Security University of China and other collaborating institutions. The current version includes two monthly snapshots, totaling 1800 Chinese telephone conversations, among which 1200 are fraud samples and 600 are nearby-domain non-fraud samples. The dataset adopts a hybrid tree-based anti-fraud generation pipeline that converts online fraud summaries into structured role-matched voice dialogues, ensuring that fraudulent and legitimate conversation paths share the same context and only diverge at critical behavioral junctures. This benchmark aims to address the limitations of existing benchmarks, which cannot sustainably update fraud patterns and lack nearby-domain negative samples, providing a more rigorous testbed for evaluating model performance in audio fraud detection under realistic ambiguous conditions.




