Southeast University Multimodal Lie Detection Dataset
收藏资源简介:
To address the lack of a Chinese context based lie detection dataset in current research, we have developed SEUMLD, which is the first publicly available multimodal lie detection dataset based on Chinese conversations. SEUMLD contains data in three modalities: video, audio, and electrocardiogram signals. In order to effectively stimulate the participants' motivation to lie, we designed a paradigm of simulated crime and simulated interrogation experiments. By recording multimodal signals of participants during simulated interrogation, SEUMLD collected data from 76 participants who had lived in a Chinese language environment for a long time, totaling 3224 conversations. This dataset provides coarse-grained annotation for identifying whether participants lie throughout the entire conversation, as well as fine-grained annotation for precise segmentation of each conversation.
针对当前研究中缺乏基于中文语境的测谎数据集这一空白,我们构建了SEUMLD——首个面向中文对话的公开多模态测谎数据集。该数据集包含三类模态数据:视频、音频与心电信号。为有效激发参与者的说谎动机,我们设计了模拟犯罪与模拟审讯实验范式。通过记录参与者在模拟审讯过程中的多模态信号,本数据集共收录76名长期处于中文语言环境下的参与者的数据,总计3224段对话。本数据集提供两类标注:其一为用于判别参与者在整段对话中是否说谎的粗粒度标注,其二为用于对单段对话进行精准分段的细粒度标注。




