| Issue |
EPJ Web Conf.
Volume 374, 2026
1st International Conference on Electronic, Optical Devices and Intelligent Systems (ICEODIS 2026)
|
|
|---|---|---|
| Article Number | 02006 | |
| Number of page(s) | 8 | |
| Section | Artificial Intelligence and Data Science | |
| DOI | https://doi.org/10.1051/epjconf/202637402006 | |
| Published online | 24 June 2026 | |
https://doi.org/10.1051/epjconf/202637402006
Hybrid CNN-DistilBERT and CNN-DistilBERT-SVM architectures for multimodal sentiment analysis on speech and text data
1
SDIA Team, UMP, Nador, Morocco
2
Laboratory of Engineering Sciences, USMBA, Taza, Morocco
3
Laboratory of Applied Mathematics and Information Systems, UMP, Nador, Morocco
Published online: 24 June 2026
Abstract
The task of sentiment analysis in human speech (acoustic and textual) is crucial for the development of empathic artificial intelligence systems in the fields of human-computer interaction (HCI) and affective computing. Accurate classification of emotional valence (positive and negative) improves the adaptability of virtual assistants and customer service platforms. The positive category includes the emotions happy and neutral, while the negative category includes angry, sad, surprised, disgust and fear. However, model performance is often compromised by speech variations caused by speaker characteristics, linguistic diversity and recording conditions. In this work, we propose a multimodal deep learning method that integrates pre-processed audio (features) with automatically generated transcribed text from Whisper. Our method combines features derived from MFCC and processed with convolutional neural networks (CNN) with semantic integrations extracted using DistilBERT, a compact and efficient Transformer-based NLP model. We develop two new hybrid architectures, CNN-DistilBERT and CNN-DistilBERT-SVM; these models are evaluated on the SAVEE, RAVDESS and TESS datasets, as well as on their combinations of datasets (TESS+SAVEE, TESS+RAVDESS, RAVDESS+SAVEE and TESS+RAVDESS+SAVEE). The CNN-DistilBERT model achieved the highest accuracy of 99.48% on the combined RAVDESS+SAVEE dataset, while the CNN-DistilBERT-SVM version achieved 98.96% accuracy on RAVDESS dataset.
© The Authors, published by EDP Sciences, 2026
This is an Open Access article distributed under the terms of the Creative Commons Attribution License 4.0, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Current usage metrics show cumulative count of Article Views (full-text article views including HTML views, PDF and ePub downloads, according to the available data) and Abstracts Views on Vision4Press platform.
Data correspond to usage on the plateform after 2015. The current usage metrics is available 48-96 hours after online publication and is updated daily on week days.
Initial download of the metrics may take a while.

