Vision-Language Modeling for Vietnamese Multimodal Fake News Classification
收藏资源简介:
We are releasing a mini subset of the research dataset for specific reasons. This dataset is designed to reflect the characteristics of online misinformation in Vietnam, where misleading content spreads not only through news-style articles but also via social media posts combining text and visual evidence. Unlike datasets directly repurposed from existing public repositories, this proposed dataset was constructed using a combination of methods: data restructuring, web crawling, social media data collection, manual supplementation, controlled synthetic data generation, and human verification. Binary Classification.• Real (0): verified information published by official or reputable news sources.• Fake (1): non-real information, including fab ricated, distorted, misleading, contextually mis matched, or unverifiable content.Four-Class Classification.• Real (0): genuine information from official or verified sources, where both textual and visual evidence are consistent with the reported event.• False Content (1): content whose main claims, textual statements, or visual elements are fabri cated, manipulated, or deliberately distorted.• False Context (2): content in which the text or image may be individually authentic, but the image is paired with a misleading caption, event, location, or time period.• Unverified (3): content consisting mainly of per sonal opinions, experiential advice, lifestyle tips, speculative statements, or belief-based claims that lack reliable verification or scientific support.



