HPSU (Human-level Perception in Spoken Speech Understanding) 和 HPSC (Human-level Perception Spoken Speech Caption)
收藏资源简介:
HPSU是由中山大学与腾讯联合构建的大规模语音理解评测基准,旨在全面评估语音大语言模型在真实场景中的人类级感知与认知能力。该数据集包含超过20,000个经过专家验证的英语和汉语口语理解样本,数据源自电影片段和社交媒体视频等开放域场景,涵盖多样化的情境与自然说话风格。其构建采用半自动标注流程,融合音频、文本和视觉信息以实现高效精准的多模态标注。该数据集主要应用于语音理解模型的深度评估与优化,致力于解决模型在潜在意图推断与隐含情感理解等方面与人类能力的差距问题。
HPSU is a large-scale speech understanding evaluation benchmark jointly developed by Sun Yat-sen University and Tencent, aiming to comprehensively assess the human-level perception and cognitive capabilities of speech-oriented large language models in real-world scenarios. This dataset contains over 20,000 expert-validated spoken language understanding samples in both English and Chinese, sourced from open-domain scenarios such as movie clips and social media videos, covering diverse contexts and natural speaking styles. Its construction adopts a semi-automatic annotation pipeline, integrating audio, text and visual information to enable efficient and accurate multi-modal annotation. This dataset is mainly used for in-depth evaluation and optimization of speech understanding models, and is committed to bridging the performance gap between models and human capabilities in aspects like latent intent inference and implicit emotion understanding.
HPSU-Benchmark 数据集概述
数据集名称
HPSU-Benchmark
核心描述
HPSU是一个面向真实世界口语语音理解的基准测试,旨在评估人类水平的感知能力。
数据集目的
为真实世界口语语音理解任务提供人类水平感知能力的评估基准。

- 1HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding中山大学、腾讯 · 2025年



