Impact Factor
Call For Paper
Volume 12 Issue 08
August 2026
Author(s)
Abstract
Speech Emotion Recognition Is An Essential Component For Applications Like Education And Human-computer Interaction [1]. While Deep Neural Networks (DNNs) Have Advanced The Field, Many Studies Ignore The Semantic Information Present Within The Speech Signal [2]. This Paper Proposes A Novel Framework Designed To Capture Both Semantic And Paralinguistic Information [5]. The Model Consists Of A Semantic Feature Extractor And A Paralinguistic Feature Extractor, Which Are Fused Together Using A Novel Attention Mechanism Into A Unified Representation. This Representation Is Then Processed By A Long Short-Term Memory (LSTM) Network To Model Temporal Dynamics [23]. Evaluated On The SEWA Dataset From The AVEC Challenge [16], The Model Achieves State-of-the-art Results In The Valence And Liking Dimensions.
Keywords
Paper ID
IJSARTV12I6105702
Publication Date
June 18, 2026
Research Area
Computer Science And Information Technology