The FETE dataset comprises facial expressions and touch gestures to fully utilize the interactive information from the two modalities during HRI. The dataset is organized into ten folders, where each
MuSe-Trust of MuSe2020: Predicting the level of trustworthiness of user-generated audio-visual content in a sequential manner utilising a diverse range of features and (optional) emotional (arousal an
CH-SIMS Dataset: This dataset consists of 2,281 video clips from different data sources, with a total of 474 speakers. The modal categories include three modalities: vision, speech, and text. The SIMS