
320
PART | III Multimodal Human-Computer and Human-to-Human Interaction
and segment labelling was evaluated at the frame-level based on
a convex combination of precision and recall (instead of using a
more standard measure similar to the Word Error Rate in speech
recognition that might not be meaningful when recognising binary
sequences). Overall, the results were promising (some of the best
reported precision/recall combinations were 63/85 and 77/60) and
indicated that combining multiple audio cues outperformed the use of
individual cues, that audio-only cues outperformed visual-only cues
and that audio-visual fusion brought benefits in some precision/recall ...