DeepSpeak-Agentic
收藏资源简介:
DeepSpeak-Agentic是由加州大学伯克利分校和斯坦福大学联合创建的大规模人机交互数据集,专注于研究具身AI代理与人类之间的实时对话动态。该数据集包含200段半结构化对话视频,总计37小时的多模态记录,涵盖了对话、专业、协作规划和创意四种场景,数据来源于143种不同的大语言模型、合成语音和视觉化身组合配置。数据采集通过定制化视频流系统实现,自动配对人类参与者和AI代理进行录制,并经过音频分离和内容审核处理。该数据集旨在支持多模态取证分析、人机交互模式研究,并为评估AI代理的逼真度和检测技术提供基准测试平台。
DeepSpeak-Agentic is a large-scale human-computer interaction dataset jointly created by the University of California, Berkeley and Stanford University, focusing on the real-time conversational dynamics between embodied AI agents and humans. This dataset contains 200 semi-structured conversational videos, totaling 37 hours of multimodal recordings, covering four scenarios: dialogue, professional tasks, collaborative planning, and creative creation. The data is sourced from 143 distinct combinations of large language models (LLMs), synthetic speech, and visual avatars. Data collection was implemented via a custom video streaming system that automatically pairs human participants with AI agents for recording, and the collected data underwent audio separation and content moderation processing. This dataset aims to support multimodal forensic analysis and research on human-computer interaction patterns, while providing a benchmark platform for evaluating the fidelity of AI agents and advancing relevant detection technologies.
- 1The DeepSpeak-Agentic Dataset加州大学伯克利分校; 斯坦福大学 · 2026年



