遇见数据集

Customer Service Voice Data Stream | High-Emotion Audio Dataset for LLM & NLP Training (US Accents)

收藏
Databricks2026-02-13 收录
官方服务:

资源简介:

We provide Real-World Customer Support Voice Data. This dataset captures the "long-tail" of human interactions: frustration, urgency, interruptions, and diverse acoustic conditions essential for building robust AI models. Why This Data is Unique? Most datasets represent the "Happy Path" (calm, polite speech). Our data solves the Class Imbalance problem by providing high-friction interactions. High Emotional Variance: Authentic anger, sarcasm, and relief — critical for training Sentiment Analysis and Empathetic AI. Real-World Acoustics: Background noise, phone line compression, and crosstalk (interruptions). Diverse Demographics: 90% unique speakers featuring a wide range of US dialects and accents. Dataset Specifications: Volume: We capture over 1,600 hours of new audio daily. Format: .wav, .flac Audio files paired with segmented, time-stamped transcriptions. Metadata: Rich labeling including Call Reason (Intent), Location, OS, Duration, and Speaker Turns. Privacy: All PII is redacted via automated and human-in-the-loop processes. Perfect For Training: LLMs & NLP: Fine-tuning models on conversational logic and intent recognition. ASR (Speech-to-Text): Improving accuracy on "messy" audio and diverse accents. Voicebots: Teaching agents to handle objections and emotional escalations. Customer Intelligence: Churn prediction and conflict resolution analytics. Data Origin: 100% sourced from real US consumers contacting customer support. Not synthetic, not scripted.

提供机构:
WiserBrand.com
搜集汇总
数据集介绍
Customer Service Voice Data Stream | High-Emotion Audio Dataset for LLM & NLP Training (US Accents) 数据集图片
背景与挑战
背景概述
该数据集收录了真实美国客户服务场景中的高情感语音交互,涵盖沮丧、打断等非理想状态,每日新增超1600小时音频,并配有丰富元数据(如呼叫原因、口音等),适用于训练情感分析、语音识别及对话机器人等AI模型。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务