TeleEgo
收藏资源简介:
TeleEgo是一个全面的全方位基准,专为自我中心视频流中的多人员、多场景、多任务和多模态长期记忆推理而设计。它反映了真实的个人助手场景,其中连续收集数小时甚至数天的自我中心视频数据,要求模型维护和推理记忆、理解和跨记忆推理。TeleEgo提供来自5个角色在4个日常场景中的全方位多样化自我中心数据、多模态注释(视频、叙述和语音转录)以及细粒度的问答基准(3个认知维度,12个子类别)。
TeleEgo is a comprehensive omnidirectional benchmark tailored for multi-person, multi-scenario, multi-task, and multimodal long-term memory reasoning in egocentric video streams. It reflects real-world personal assistant scenarios, where egocentric video data is continuously collected over hours or even days, requiring models to maintain, reason over memories, as well as perform comprehension and cross-memory reasoning. TeleEgo provides omnidirectional and diverse egocentric data from 5 characters across 4 daily scenarios, multimodal annotations including video, narration, and speech transcripts, and a fine-grained question-answering benchmark covering 3 cognitive dimensions and 12 subcategories.
TeleEgo 数据集概述
数据集简介
TeleEgo 是一个全面的全方位基准测试,专为以自我为中心的视频流中的多人、多场景、多任务和多模态长期记忆推理而设计。该基准测试反映了真实的个人助手场景,其中连续的自中心视频数据在数小时甚至数天内收集,要求模型维护和推理记忆、理解以及跨记忆推理。
数据集特点
- 全方位覆盖:涵盖角色、场景、任务、模态和记忆视野的全谱系
- 多模态数据:视频、叙述和语音转录
- 细粒度问答基准:3个认知维度,12个子类别
数据集规模
- 参与者:5人(性别平衡)
- 场景:
- 工作与学习
- 生活方式与日常
- 社交活动
- 外出与文化
- 录制时长:每人3天(约14.4小时/人)
- 模态:
- 以自我为中心的视频流
- 语音和对话
- 叙述和事件描述
基准测试任务
TeleEgo-QA 沿三个主要维度评估模型:
记忆
- 短期/长期/超长期记忆
- 实体追踪
- 时间比较与间隔
理解
- 因果理解
- 意图推断
- 多步推理
- 跨模态理解
跨记忆推理
- 跨时间因果关系
- 跨实体关系
- 时间链理解
每个问答实例包括:
- 问题类型:单选、多选、二元、开放性问题
数据集访问
由于隐私和许可限制,请在此处请求访问: https://huggingface.co/datasets/David0219/TeleEgo
引用
bibtex @misc{yan2025teleegobenchmarkingegocentricai, title={TeleEgo: Benchmarking Egocentric AI Assistants in the Wild}, author={Jiaqi Yan and Ruilong Ren and Jingren Liu and Shuning Xu and Ling Wang and Yiheng Wang and Yun Wang and Long Zhang and Xiangyu Chen and Changzhi Sun and Jixiang Luo and Dell Zhang and Hao Sun and Chi Zhang and Xuelong Li}, year={2025}, eprint={2510.23981}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2510.23981}, }
许可证
本项目采用 MIT 许可证。 数据集使用受限于仅研究用途许可证。




