虚拟说话人形象标准图片集及其对应的数字人模型数据集
收藏资源简介:
构建的基于多模态智能交互的奥林匹克云展厅系统集成了多模态虚拟形象交互技术,其中个性化虚拟形象生成使用单张照片进行人脸三维重建,用户仅需上传一张正面照片即可完成个性化虚拟形象构建,支持语音驱动的说话唇形动画生成。支持选定云展厅中预设的形象照片或者上传用户自身形象的照片,基于此利用基于卷积神经网络的三维人脸重建模型,通过输入照片预测人脸的参数化BFM(Basel Face Model)系数(包括人脸的身份参数、表情参数等)。此后通过冬奥会云展厅预设的光照、人脸姿势、人脸纹理等配置,结合上述BFM系数即可在云展厅中生成符合用户需求的虚拟形象。
The constructed Olympic cloud exhibition hall system based on multimodal intelligent interaction integrates multimodal virtual avatar interaction technology. For personalized virtual avatar generation, it adopts single-photo-based 3D face reconstruction: users can complete the creation of personalized virtual avatars by only uploading a single frontal photo, and the system supports voice-driven lip-sync animation generation for speech. The system allows users to either select preset avatar photos from the cloud exhibition hall or upload their own portrait photos. Based on the input photos, the Convolutional Neural Network (CNN)-based 3D face reconstruction model predicts the parametric Basel Face Model (BFM) coefficients of the face, including face identity parameters, expression parameters and other relevant parameters. Subsequently, by combining the aforementioned BFM coefficients with the preset configurations such as lighting, face pose and face texture in the Winter Olympics cloud exhibition hall, the system can generate virtual avatars that meet users' requirements in the cloud exhibition hall.




