Real-Time Multi-Modal Emotion & Stress Detector from Speech & Facial Video
Fusion model combining facial Action Units (FACS) with voice pitch and jitter to classify acute workplace stress levels.
Project Overview
Captures synchronized webcam video and microphone audio during remote work or clinical interviews. Fuses facial micro-expression classifications with acoustic speech prosody features in a late-fusion transformer to assess cognitive overload and mental exhaustion.
Captures synchronized webcam video and microphone audio during remote work or clinical interviews. Fuses facial micro-expression classifications with acoustic speech prosody features in a late-fusion transformer to assess cognitive overload and mental exhaustion.
Fusion model combining facial Action Units (FACS) with voice pitch and jitter to classify acute workplace stress levels.
Core Project Objectives
Capture continuous analog/digital sensor readings with robust noise filtering and hardware calibration.