Vui lòng dùng định danh này để trích dẫn hoặc liên kết đến tài liệu này: https://elib.vku.udn.vn/handle/123456789/7719
Nhan đề: Toward Robust Speech Emotion Recognition: A Hybrid EMix-Based Augmentation Approach
Tác giả: Dang, An
Le, Dinh Nguyen
Ha, Minh Tan
Vu, Quang Duc
Từ khoá: SER
Data augmentation
Emix
Năm xuất bản: thá-2026
Nhà xuất bản: Springer Nature
Tóm tắt: Recently, various deep learning (DL) models have been introduced to enhance speech emotion recognition (SER) accuracy. However, the scarcity and limited scale of SER datasets due to the complexity and cost of data collection often result in model overfitting, thereby restricting overall performance. In this paper, we propose a hybrid data augmentation framework that integrates the previously introduced emotion-aware EMix method with complementary time-frequency perturbation techniques to enhance both model robustness and generalization. To validate the proposed approach, we develop a deep convolutional neural network consisting of two main components: a multi-branch stem block and a DenseNet-based backbone. The stem block extracts informative features across various time-frequency scales, while the backbone offers robust representational capacity based on a pre-trained image classification model. Experimental results on two publicly available benchmark datasets confirm the effectiveness of the proposed method, achieving state-of-the-art accuracies of 82.39% on CREMA-D and 78.25% on IEMOCAP.
Mô tả: Intelligence of Things: Technologies and Applications (ICIT 2025); pp: 19-30
Định danh: https://doi.org/10.1007/978-3-032-13254-3_2
https://elib.vku.udn.vn/handle/123456789/7719
ISSN: 978-3-032-13253-6 (p)
978-3-032-13254-3 (p)
Bộ sưu tập: NĂM 2026

Các tập tin trong tài liệu này:

 Đăng nhập để xem toàn văn



Khi sử dụng các tài liệu trong Thư viện số phải tuân thủ Luật bản quyền.