遇见数据集

Corpus Data for: \"Hearing lips: on the dominance of vision in immersive cocktail party phenomena\"

收藏
DataONE2017-01-26 更新2024-06-26 收录
官方服务:

资源简介:

Immersive stereoscopic footage of a Coordinate Response Measure (CRM) recorded from two actors. The audio-visual recorded corpus consists of 8 CALLs and 32 COMMANDs per actor. The CALLs and COMMANDs are to be combined at rendering time into full sentences that always follow the same structure: “ready CALL go to COMMAND now ”. The COMMANDs consists of one in four colors (blue, green, red or white) followed by one in eight numbers (1 to 8). This generates a full combinatorial of 256 individual sentences when combined with one of the 8 CALLS (arrow, baron, charlie, eagle, hopper, laker, ringo, tiger). Additionally the dataset also includes the UV positions to texturize the semi-spheres at the rendering time. These have been calculated from the intrinsic and extrinsic calibration parameters of the cameras to facilitate the correct rendering of the video footage. Our system for recording the actors consists of a custom wide-angle stereo camera system made of two Grasshopper 3 cameras with fisheye Fujinon lenses (2.7mm focal length) reaching 185 degrees of Field of View (FoV). The cameras were mounted parallel to each other and separated by 65 mm distance (average human interpupillary distance39) to provide stereoscopic capturing. The video is encoded in H264 format reaching 28-30 frames per second encoding speed at 1600x1080 resolution per camera/eye. The audio was recorded through a near range microphone at a 44kHz sampling rate and 99kbps and both the audio and video are synchronized within 10ms range and saved in mp4 format. The recording room was equipped for professional recording with monobloc LED lighting and chromakey screen. The actor sat at 1 meter distance from the camera recording setup and read the corpus sentences when presented on the screen behind the cameras. The actors were recorded separately in two sessions, seating each at 30 degrees from the bisection, and their videos can be synthetically attached at the rendering time. In the post processing the audio was equalized for all words, and the video was stitched to combine the actors and generate the full the corpus. Sentences were band passed at 80Hz to 16kHz. The corpus sentences are temporally aligned within the range of 64ms in our case, which is below the described 200ms to be perceived. So two or more CRMs can be played synchronously generating an overlap.

本数据集包含两名演员录制的坐标响应度量(Coordinate Response Measure, CRM)沉浸式立体影像。该音视频语料库中,每名演员对应8条呼叫指令(CALL)与32条命令(COMMAND)。在渲染阶段,需将呼叫指令与命令组合为固定结构的完整句子:"ready CALL go to COMMAND now"。命令由四类颜色(蓝、绿、红、白)之一与八个数字(1至8)之一组合而成,结合8条呼叫指令(arrow、baron、charlie、eagle、hopper、laker、ringo、tiger),可生成共计256条独立的完整句子。此外,本数据集还包含用于渲染阶段对半球模型进行纹理映射的UV坐标,该坐标基于相机的内外参校准参数计算得到,可保障影像渲染的正确性。本次录制采用的定制化广角立体相机系统由两台搭载富士能(Fujinon)鱼眼镜头(焦距2.7mm,视场角(Field of View, FoV)185°)的草蜢3(Grasshopper 3)相机组成。两台相机平行安装,间距为65mm(接近人类平均瞳孔间距39),以实现立体拍摄。视频采用H264格式编码,单相机/单眼分辨率为1600×1080,编码帧率达28-30帧每秒。音频通过近距离麦克风录制,采样率为44kHz,码率99kbps。音视频同步误差控制在10ms以内,最终以MP4格式存储。录制场地为专业录音棚,配备一体化LED照明系统与抠像绿幕。演员距相机录制系统1米就座,根据相机后方屏幕显示的语料句子进行朗读。两名演员分两次单独录制,各自与相机中轴线呈30°夹角,其影像可在渲染阶段进行合成拼接。后期处理环节中,所有语音均经过均衡化处理,视频经拼接合成两名演员的影像以生成完整语料库。语音信号经过80Hz至16kHz的带通滤波。本数据集中语料句子的时间对齐误差控制在64ms以内,低于200ms的可感知阈值,因此可同步播放两条或多条CRM以实现音频叠加效果。

创建时间:
2023-11-21
二维码
社区交流群
二维码
科研交流群
商业服务