LIA: LLM-based Information distillation and Activity recognition
收藏资源简介:
This repository contains the data and supporting materials associated with the study: “LIA: LLM-based Information Distillation and Activity Recognition – A Training-Free Framework for Zero-Shot Temporal Action Segmentation.” Overview In this study, we introduce LIA (LLM-based Information Distillation and Activity Recognition), a training-free framework for zero-shot temporal action segmentation. The method leverages the frozen multimodal reasoning capabilities of pretrained Large Language Models (LLMs) and introduces a multistep semantic distillation pipeline that converts noisy real-world narrations into a structured taxonomy of action primitives without parameter updates or domain-specific fine-tuning. The framework integrates Bayesian temporal inference with Visual Question Answering (VQA)-based reasoning to enforce temporal coherence in untrimmed videos. Evaluation was conducted on the YouTube Instructions (YTI) dataset.



