Modeling naturalistic affective states via facial and vocal expressions recognition

George Caridakis*, Lori Malatesta, Loic Kessous, Noam Amir, Amaryllis Raouzaiou, Kostas Karpouzis

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review


Affective and human-centered computing are two areas related to HCI which have attracted attention during the past years. One of the reasons that this may be attributed to, is the plethora of devices able to record and process multimodal input from the part of the users and adapt their functionality to their preferences or individual habits, thus enhancing usability and becoming attractive to users less accustomed with conventional interfaces. In the quest to receive feedback from the users in an unobtrusive manner, the visual and auditory modalities allow us to infer the users' emotional state, combining information both from facial expression recognition and speech prosody feature extraction. In this paper, we describe a multi-cue, dynamic approach in naturalistic video sequences. Contrary to strictly controlled recording conditions of audiovisual material, the current research focuses on sequences taken from nearly real world situations. Recognition is performed via a 'Simple Recurrent Network' which lends itself well to modeling dynamic events in both user's facial expressions and speech. Moreover this approach differs from existing work in that it models user expressivity using a dimensional representation of activation and valence, instead of detecting the usual 'universal emotions' which are scarce in everyday human-machine interaction. The algorithm is deployed on an audiovisual database which was recorded simulating human-human discourse and, therefore, contains less extreme expressivity and subtle variations of a number of emotion labels.

Original languageEnglish
Title of host publicationICMI'06
Subtitle of host publication8th International Conference on Multimodal Interfaces, Conference Proceedings
Number of pages9
StatePublished - 2006
Externally publishedYes
EventICMI'06: 8th International Conference on Multimodal Interfaces - Banff, AB, Canada
Duration: 2 Nov 20064 Nov 2006

Publication series

NameICMI'06: 8th International Conference on Multimodal Interfaces, Conference Proceeding


ConferenceICMI'06: 8th International Conference on Multimodal Interfaces
CityBanff, AB


  • Affective interaction
  • Facial expression recognition
  • Image processing
  • Multimodal analysis
  • Naturalistic data
  • Prosodic feature extraction
  • User modeling


Dive into the research topics of 'Modeling naturalistic affective states via facial and vocal expressions recognition'. Together they form a unique fingerprint.

Cite this