TeleAntiFraud-28k
收藏资源简介:
TeleAntiFraud-28k是一个由中国人民移动互联网有限公司创建的开源音频文本慢思考数据集,专为电信欺诈检测设计。该数据集通过三种策略构建:使用自动语音识别技术生成隐私保护的文本样本,利用大型语言模型进行语义增强,以及通过多智能体对抗框架模拟新型欺诈策略。数据集包含28511个经过严格处理的语音文本对,并提供了详细的欺诈推理注释。数据集分为三种任务:场景分类、欺诈检测、欺诈类型分类,并为研究者提供了一个统一评估平台TeleAntiFraud-Bench,用于评估不同模型在电信欺诈检测任务中的性能。
TeleAntiFraud-28k is an open-source audio-text slow-thinking dataset created by China People's Mobile Internet Co., Ltd., specifically tailored for telecom fraud detection. This dataset is constructed via three strategies: 1) generating privacy-preserving text samples using automatic speech recognition (ASR) technology, 2) performing semantic enhancement with large language models (LLMs), and 3) simulating novel fraud tactics through a multi-agent adversarial framework. It comprises 28,511 rigorously processed audio-text pairs, accompanied by detailed fraud reasoning annotations. The dataset encompasses three task types: scenario classification, fraud detection, and fraud type classification, and provides researchers with a unified evaluation platform named TeleAntiFraud-Bench for evaluating the performance of various models on telecom fraud detection tasks.

- 1TeleAntiFraud-28k: A Audio-Text Slow-Thinking Dataset for Telecom Fraud Detection中国移动互联网有限公司 · 2025年



