遇见数据集

SocialOmni

收藏
魔搭社区2026-09-06 更新2026-09-06 收录
官方服务:

资源简介:

# SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models [**Paper**](https://huggingface.co/papers/2603.16859) | [**GitHub**](https://github.com/MAC-AutoML/SocialOmni) | [**Project Page**](https://huggingface.co/datasets/alexisty/SocialOmni) SocialOmni is a comprehensive benchmark designed to evaluate the **audio-visual social interactivity** of Omni-modal Large Language Models (OLMs). Instead of focusing solely on static accuracy, SocialOmni measures whether a model can behave appropriately in real dialogues by evaluating three tightly coupled dimensions: - **Who** is speaking: speaker separation and identification. - **When** to enter: interruption timing control. - **How** to respond: natural interruption generation. The benchmark features 2,000 perception samples and a quality-controlled diagnostic set of 209 interaction-generation instances with strict temporal and contextual constraints, complemented by controlled audio-visual inconsistency scenarios to test model robustness. ## Tasks ### Task I: Perception (`who`) Given a video clip and a timestamp `t`, the model identifies the active speaker from a set of candidates (e.g., A, B, C, or D). ### Task II: Interaction Generation (`when` + `how`) Given a video prefix and a candidate speaker, the model performs two sub-tasks: - **Q1 (`when`)**: Decide if the speaker should interrupt immediately. - **Q2 (`how`)**: If an interruption is appropriate, generate the natural and contextually coherent content for that interruption. ## Citation ```bibtex @article{xie2026socialomni, title={SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models}, author={Tianyu Xie and Jinfa Huang and Yuexiao Ma and Rongfang Luo and Yan Yang and Wang Chen and Yuhui Zeng and Ruize Fang and Yixuan Zou and Xiawu Zheng and Jiebo Luo and Rongrong Ji}, journal={arXiv preprint arXiv:2603.16859}, year={2026} } ```

提供机构:
maas
创建时间:
2026-09-02
二维码
社区交流群
二维码
科研交流群
商业服务