Please use this identifier to cite or link to this item: https://elib.vku.udn.vn/handle/123456789/7719
Title: Toward Robust Speech Emotion Recognition: A Hybrid EMix-Based Augmentation Approach
Authors: Dang, An
Le, Dinh Nguyen
Ha, Minh Tan
Vu, Quang Duc
Keywords: SER
Data augmentation
Emix
Issue Date: May-2026
Publisher: Springer Nature
Abstract: Recently, various deep learning (DL) models have been introduced to enhance speech emotion recognition (SER) accuracy. However, the scarcity and limited scale of SER datasets due to the complexity and cost of data collection often result in model overfitting, thereby restricting overall performance. In this paper, we propose a hybrid data augmentation framework that integrates the previously introduced emotion-aware EMix method with complementary time-frequency perturbation techniques to enhance both model robustness and generalization. To validate the proposed approach, we develop a deep convolutional neural network consisting of two main components: a multi-branch stem block and a DenseNet-based backbone. The stem block extracts informative features across various time-frequency scales, while the backbone offers robust representational capacity based on a pre-trained image classification model. Experimental results on two publicly available benchmark datasets confirm the effectiveness of the proposed method, achieving state-of-the-art accuracies of 82.39% on CREMA-D and 78.25% on IEMOCAP.
Description: Intelligence of Things: Technologies and Applications (ICIT 2025); pp: 19-30
URI: https://doi.org/10.1007/978-3-032-13254-3_2
https://elib.vku.udn.vn/handle/123456789/7719
ISSN: 978-3-032-13253-6 (p)
978-3-032-13254-3 (p)
Appears in Collections:NĂM 2026

Files in This Item:

 Sign in to read



Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.