Neural Face Models for Example-Based Visual Speech Synthesis

Paier, Wolfgang; Hilsmann, Anna; Eisert, Peter

Full-text links:

Download:

Current browse context:

cs.CV

< prev | next >

new | recent | 2009

Change to browse by:

Computer Science > Computer Vision and Pattern Recognition

Title: Neural Face Models for Example-Based Visual Speech Synthesis

Authors: Wolfgang Paier, Anna Hilsmann, Peter Eisert

(Submitted on 22 Sep 2020)

Abstract: Creating realistic animations of human faces with computer graphic models is still a challenging task. It is often solved either with tedious manual work or motion capture based techniques that require specialised and costly hardware. Example based animation approaches circumvent these problems by re-using captured data of real people. This data is split into short motion samples that can be looped or concatenated in order to create novel motion sequences. The obvious advantages of this approach are the simplicity of use and the high realism, since the data exhibits only real deformations. Rather than tuning weights of a complex face rig, the animation task is performed on a higher level by arranging typical motion samples in a way such that the desired facial performance is achieved. Two difficulties with example based approaches, however, are high memory requirements as well as the creation of artefact-free and realistic transitions between motion samples. We solve these problems by combining the realism and simplicity of example-based animations with the advantages of neural face models. Our neural face model is capable of synthesising high quality 3D face geometry and texture according to a compact latent parameter vector. This latent representation reduces memory requirements by a factor of 100 and helps creating seamless transitions between concatenated motion samples. In this paper, we present a marker-less approach for facial motion capture based on multi-view video. Based on the captured data, we learn a neural representation of facial expressions, which is used to seamlessly concatenate facial performances during the animation procedure. We demonstrate the effectiveness of our approach by synthesising mouthings for Swiss-German sign language based on viseme query sequences.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
ACM classes:	I.3.5; I.4.8
Cite as:	arXiv:2009.10361 [cs.CV]
	(or arXiv:2009.10361v1 [cs.CV] for this version)

Submission history

From: Wolfgang Paier [view email]
[v1] Tue, 22 Sep 2020 07:35:33 GMT (6242kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2009.10361

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computer Vision and Pattern Recognition

Title: Neural Face Models for Example-Based Visual Speech Synthesis

Submission history