Please use this identifier to cite or link to this item:
https://elib.vku.udn.vn/handle/123456789/7719| Title: | Toward Robust Speech Emotion Recognition: A Hybrid EMix-Based Augmentation Approach |
| Authors: | Dang, An Le, Dinh Nguyen Ha, Minh Tan Vu, Quang Duc |
| Keywords: | SER Data augmentation Emix |
| Issue Date: | May-2026 |
| Publisher: | Springer Nature |
| Abstract: | Recently, various deep learning (DL) models have been introduced to enhance speech emotion recognition (SER) accuracy. However, the scarcity and limited scale of SER datasets due to the complexity and cost of data collection often result in model overfitting, thereby restricting overall performance. In this paper, we propose a hybrid data augmentation framework that integrates the previously introduced emotion-aware EMix method with complementary time-frequency perturbation techniques to enhance both model robustness and generalization. To validate the proposed approach, we develop a deep convolutional neural network consisting of two main components: a multi-branch stem block and a DenseNet-based backbone. The stem block extracts informative features across various time-frequency scales, while the backbone offers robust representational capacity based on a pre-trained image classification model. Experimental results on two publicly available benchmark datasets confirm the effectiveness of the proposed method, achieving state-of-the-art accuracies of 82.39% on CREMA-D and 78.25% on IEMOCAP. |
| Description: | Intelligence of Things: Technologies and Applications (ICIT 2025); pp: 19-30 |
| URI: | https://doi.org/10.1007/978-3-032-13254-3_2 https://elib.vku.udn.vn/handle/123456789/7719 |
| ISSN: | 978-3-032-13253-6 (p) 978-3-032-13254-3 (p) |
| Appears in Collections: | NĂM 2026 |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.