CH-SIMS Dataset: This dataset consists of 2,281 video clips from different data sources, with a total of 474 speakers. The modal categories include three modalities: vision, speech, and text. The SIMS
MuSe-Trust of MuSe2020: Predicting the level of trustworthiness of user-generated audio-visual content in a sequential manner utilising a diverse range of features and (optional) emotional (arousal an
[Post-challenge] MuSe-Topic of MuSe2020: Predicting 10-class domain-specific topics as the target of 3-class (low, medium, high) emotions of valence and arousal. This package includes only MuSe-Topic