
274
PART | III Multimodal Human–Computer and Human-to-Human Interaction
There have been attempts to use the preprocessing step of one
TTS engine and apply the obtained timing to another TTS engine;
however, no results have been published yet. A promising approach
is to use a machine learning technique to predict the timing. This has
been done in [54] using neural networks with acceptable results for
gesture synchronisation, though not for coarticulation.
13.8 CONCLUSION
This chapter provides a structured overview of issues involved in
generating multimodal output consisting of speech, facial motion
and gestures for the purpose of HCI. We concentrated