Vui lòng dùng định danh này để trích dẫn hoặc liên kết đến tài liệu này:
https://elib.vku.udn.vn/handle/123456789/7719| Nhan đề: | Toward Robust Speech Emotion Recognition: A Hybrid EMix-Based Augmentation Approach |
| Tác giả: | Dang, An Le, Dinh Nguyen Ha, Minh Tan Vu, Quang Duc |
| Từ khoá: | SER Data augmentation Emix |
| Năm xuất bản: | thá-2026 |
| Nhà xuất bản: | Springer Nature |
| Tóm tắt: | Recently, various deep learning (DL) models have been introduced to enhance speech emotion recognition (SER) accuracy. However, the scarcity and limited scale of SER datasets due to the complexity and cost of data collection often result in model overfitting, thereby restricting overall performance. In this paper, we propose a hybrid data augmentation framework that integrates the previously introduced emotion-aware EMix method with complementary time-frequency perturbation techniques to enhance both model robustness and generalization. To validate the proposed approach, we develop a deep convolutional neural network consisting of two main components: a multi-branch stem block and a DenseNet-based backbone. The stem block extracts informative features across various time-frequency scales, while the backbone offers robust representational capacity based on a pre-trained image classification model. Experimental results on two publicly available benchmark datasets confirm the effectiveness of the proposed method, achieving state-of-the-art accuracies of 82.39% on CREMA-D and 78.25% on IEMOCAP. |
| Mô tả: | Intelligence of Things: Technologies and Applications (ICIT 2025); pp: 19-30 |
| Định danh: | https://doi.org/10.1007/978-3-032-13254-3_2 https://elib.vku.udn.vn/handle/123456789/7719 |
| ISSN: | 978-3-032-13253-6 (p) 978-3-032-13254-3 (p) |
| Bộ sưu tập: | NĂM 2026 |
Khi sử dụng các tài liệu trong Thư viện số phải tuân thủ Luật bản quyền.