This open-source dataset consists of 4.25 hours of transcribed Guangzhou Cantonese conversational speech on certain topics, where ten conversations between ten pairs of speakers were contained.
The MODALITY corpus is one of the multimodal database of word recordings in English. It consists of over 30 hours of multimodal recordings. The database contains high-resolution, high-framerate stereo