Towards Efficient Inference in Video Understanding
收藏数据链接:
官方服务:
资源简介:
While Vision Transformers and Large Language Models have advanced video understanding, their computational demands have pose significant challenges. This thesis explores the temporal redundancy issue in video understanding tasks, developing efficient video understanding models that improve inference efficiency while maintaining competitive model performance for real-world applications.
创建时间:
2025-05-22




