karanverma19/trilingual_fraud_consumer_protection_final_v3
收藏资源简介:
该数据集专注于英语、印地语和旁遮普语的真实世界欺诈检测和消费者保护场景,特别是在移民和就业背景下。V3版本的关键改进包括为每个标签添加了推理、包含行为欺诈信号(如紧迫感、支付请求、权威滥用)、提供可操作指导(用户应采取的措施)以及反映真实世界的代码混合通信模式。数据集的重要性在于它不仅关注分类,还模拟决策制定和用户安全指导,使其对现实世界的AI系统更有用。每个条目包括用户消息(多语言、代码混合)和助手响应(分类、推理、行动)。使用案例包括欺诈检测系统、AI安全和对齐、消费者保护工具以及多语言NLP研究。该数据集与Uncharted Data Challenge对齐,涉及服务不足的领域(移民欺诈)、代表性不足的语言(旁遮普语、印地语)以及现实世界的影响(保护弱势用户)。
This dataset focuses on real-world fraud detection and consumer protection scenarios in English, Hindi, and Punjabi, particularly in immigration and employment contexts. Key improvements in V3 include adding reasoning to each label, incorporating behavioral fraud signals (urgency, payment request, authority misuse), providing actionable guidance (what user should do), and reflecting real-world code-mixed communication patterns. The dataset matters because it goes beyond classification to model decision-making and user safety guidance, making it more useful for real-world AI systems. Each entry includes user message (multilingual, code-mixed) and assistant response (classification, reasoning, action). Use cases include fraud detection systems, AI safety and alignment, consumer protection tools, and multilingual NLP research. The dataset aligns with the Uncharted Data Challenge by addressing an underserved domain (immigration fraud), underrepresented languages (Punjabi, Hindi), and real-world impact (protecting vulnerable users).




