GD-ML/Taming-Hallucinations
收藏资源简介:
该存储库托管了“DualityVidQA”,这是由论文《Taming Hallucinations: Boosting MLLMs Video Understanding via Counterfactual Video Generation》引入的大规模配对视频-QA数据集。Taming Hallucinations 引入了 DualityForge,一个基于可控扩散的框架,将真实视频转换为反事实视频,自动生成配对(真实/反事实)视频及其问答数据以进行对比训练。基于 DualityVidQA 和提出的 DNA-Train SFT-RL 机制(使用 ℓ1 归一化优势),该方法将多模态 LLM 的幻觉减少了 24%,并在多个基准测试中表现出强大的泛化能力。
This repository hosts DualityVidQA, the large-scale paired video–QA dataset introduced in Taming Hallucinations: Boosting MLLMs Video Understanding via Counterfactual Video Generation. Taming Hallucinations introduces DualityForge, a controllable diffusion-based framework that turns real videos into counterfactual ones, automatically generating paired (real / counterfactual) videos together with their question–answer data for contrastive training. Built on top of DualityVidQA and the proposed DNA-Train SFT–RL regime with ℓ1-normalized advantages, our approach reduces hallucinations in multimodal LLMs by 24% and shows strong generalization across benchmarks.




