We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

eess.AS

Change to browse by:

References & Citations

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Multimodal generation of upper-facial and head gestures with a Transformer Network using speech and text

Abstract: We propose a semantically-aware speech driven method to generate expressive and natural upper-facial and head motion for Embodied Conversational Agents (ECA). In this work, we tackle two key challenges: produce natural and continuous head motion and upper-facial gestures. We propose a model that generates gestures based on multimodal input features: the first modality is text, and the second one is speech prosody. Our model makes use of Transformers and Convolutions to map the multimodal features that correspond to an utterance to continuous eyebrows and head gestures. We conduct subjective and objective evaluations to validate our approach.
Subjects: Audio and Speech Processing (eess.AS)
Cite as: arXiv:2110.04527 [eess.AS]
  (or arXiv:2110.04527v1 [eess.AS] for this version)

Submission history

From: Nicolas Obin [view email]
[v1] Sat, 9 Oct 2021 09:38:40 GMT (954kb,D)
[v2] Sat, 21 May 2022 10:38:33 GMT (2115kb,D)

Link back to: arXiv, form interface, contact.