
258
PART | III Multimodal Human–Computer and Human-to-Human Interaction
gestures. We start by introducing a basic audio-visual (AV) speech
synthesis system that generates simple lip motion from input text
using a text-to-speech (TTS, see Chapter 3) engine and an animation
system. Throughout this chapter, we gradually extend and improve
this system first with coarticulation, then full facial motion and ges-
tures and finally we present it in the context of a full embodied
conversational agent (ECA) system. At each level, we present key
concepts and discuss existing systems.
We concentrate on real-time interactive systems, as necessary for
human–computer interaction ...