遇见数据集

Towards Efficient Inference in Video Understanding

收藏
Monash University Figshare2026-02-11 更新2026-07-03 收录
官方服务:

资源简介:

While Vision Transformers and Large Language Models have advanced video understanding, their computational demands have pose significant challenges. This thesis explores the temporal redundancy issue in video understanding tasks, developing efficient video understanding models that improve inference efficiency while maintaining competitive model performance for real-world applications.

创建时间:
2025-05-22
二维码
社区交流群
二维码
科研交流群
商业服务