Journal of Multimedia Technology & Recent Advancements Review Article
Advancements in AI-Driven Sound Spectrogram Analysis: From Deep Learning to Quantum and Neuromorphic Processing
Abstract
The rapid advancement of artificial intelligence (AI) has significantly reshaped the field of audio signal processing, with sound spectrogram analysis emerging as a central research focus. Spectrograms provide a rich time–frequency representation of audio signals, making them particularly suitable for data-driven learning approaches. This paper presents an in-depth and original review of modern AI-based techniques applied to spectrogram analysis, highlighting their growing impact across critical application areas such as healthcare diagnostics, security systems, environmental surveillance, and intelligent multimedia processing. We examine a range of state-of-the-art methodologies, including convolutional neural networks (CNNs), transformer-based models, and adaptive spectral estimation strategies, emphasizing how each approach leverages spectro-temporal patterns for improved feature extraction and classification. Through multiple case studies – covering deepfake audio detection, disease diagnosis from biomedical sounds, and acoustic scene classification – we demonstrate that AI-driven models consistently outperform conventional signal processing methods in terms of accuracy, generalization, and noise robustness. A key contribution of this work is the discussion of explainable artificial intelligence (XAI) techniques, which enhance model transparency and trust by revealing how spectral features influence decision-making. Experimental evaluations show that transformer architectures equipped with spectrogram-aware attention mechanisms achieve up to 92.4% accuracy on standard benchmark datasets, exceeding CNN-based models by 6.8%. Finally, the paper identifies current technical challenges and outlines future research directions, including multimodal data fusion, efficient edge deployment, and the potential role of quantum-enhanced spectral analysis in next-generation audio intelligence systems.
Keywords
References (30)
- Anagha R, Arya A, Narayan VH, Abhishek S, Anjali T. Audio Deepfake Detection Using Deep Learning. 2023 12th International Conference on System Modeling & Advancement in Research Trends (SMART). 2023:176-181. doi:10.1109/smart59791.2023.10428163
- Balamurugan A, Teo SG, Yang J, Peng Z, Xulei Y, Zeng Z. ResHNet: Spectrograms Based Efficient Heart Sounds Classification Using Stacked Residual Networks. 2019 IEEE EMBS International Conference on Biomedical & Health Informatics (BHI). 2019:1-4. doi:10.1109/bhi.2019.8834578
- Nogueira AFR, Oliveira HS, Machado JJM, Tavares JMRS. Transformers for Urban Sound Classification—A Comprehensive Performance Evaluation. Sensors. 2022;22(22):8874. doi:10.3390/s22228874
- Jones CC, Gannon ZE, Blunt SD, Allen CT, Martone AF. An Adaptive Spectrogram Estimator to Enhance Signal Characterization. 2022 IEEE Radar Conference (RadarConf22). 2022:1-6. doi:10.1109/radarconf2248738.2022.9764186
- Tsui BMW, Xu J, Rittenbach A, Chen S, El-Sharkaway AM, Edelstein WA, et al. High performance SPECT system for simultaneous SPECT-MR imaging of small animals. 2011 IEEE Nuclear Science Symposium Conference Record. 2011:3178-3182. doi:10.1109/nssmic.2011.6153652
- Tian B, Pang Y, Huzaifa M, Wang S, Adve S. Towards Energy-Efficiency by Navigating the Trilemma of Energy, Latency, and Accuracy. 2024 IEEE International Symposium on Mixed and Augmented Reality (ISMAR). 2024:913-922. doi:10.1109/ismar62088.2024.00107
- Xia Y, Zhao Z. Cross-modal Background Suppression for Audio-Visual Event Localization. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022:19957-19966. doi:10.1109/cvpr52688.2022.01936
- Abdul ZK, Al-Talabani AK. Mel Frequency Cepstral Coefficient and its Applications: A Review. IEEE Access. 2022;10:122136-122158. doi:10.1109/access.2022.3223444
- Cerezuela-Escudero E, Jimenez-Fernandez A, Paz-Vicente R, Dominguez-Morales M, Linares-Barranco A, Jimenez-Moreno G. Musical notes classification with neuromorphic auditory system using FPGA and a convolutional spiking network. 2015 International Joint Conference on Neural Networks (IJCNN). 2015:1-7. doi:10.1109/ijcnn.2015.7280619
- Li F, Zhang Z, Wang L, Liu W. Heart sound classification based on improved mel-frequency spectral coefficients and deep residual learning. Frontiers in Physiology. 2022;13. doi:10.3389/fphys.2022.1084420
- Kim G, Han DK, Ko H. SpecMix: a mixed sample data augmentation method for training with time-frequency domain features. [preprint]. 2021. arXiv:2108.03020. doi:10.48550/arXiv.2108.03020.
- Foresti GL, Regazzoni CS. Multisensor data fusion for autonomous vehicle navigation in risky environments. IEEE Transactions on Vehicular Technology. 2002;51(5):1165-1185. doi:10.1109/tvt.2002.800629
- Wang H, Zou Y, Wang W. SpecAugment++: a hidden space data augmentation method for acoustic scene classification. [preprint]. 2021. arXiv:2103.16858. doi:10.48550/arXiv.2103.16858.
- Han J, Matuszewski M, Sikorski O, Sung H, Cho H. Randmasking Augment: A Simple and Randomized Data Augmentation For Acoustic Scene Classification. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2023:1-5. doi:10.1109/icassp49357.2023.10095001
- Thuillier E, Gamper H, Tashev IJ. Spatial Audio Feature Discovery with Convolutional Neural Networks. 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2018:6797-6801. doi:10.1109/icassp.2018.8462315
- Schluter J, Gutenbrunner G. EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use. 2022 30th European Signal Processing Conference (EUSIPCO). 2022:205-208. doi:10.23919/eusipco55093.2022.9909910
- Nguyen TTM, Nguyen DD, Luong CM. Vietnamese Speaker Verification With Mel-Scale Filter Bank Energies and Deep Learning. IEEE Access. 2024;12:150114-150122. doi:10.1109/access.2024.3479092
- Wang J, Li J, Tan X. Spectral-Spatial Symmetrical Aggregation Cross-Linking Multi-Modal Data Fusion Network. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2022:5098-5102. doi:10.1109/icassp43922.2022.9747570
- Barahona S, de Benito-Gorrón D, Toledano DT, Ramos D. Enhancing Conformer-Based Sound Event Detection Using Frequency Dynamic Convolutions and BEATs Audio Embeddings. IEEE/ACM Transactions on Audio, Speech, and Language Processing. 2024;32:3896-3907. doi:10.1109/taslp.2024.3444490
- Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, et al. A Survey on Vision Transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2023;45(1):87-110. doi:10.1109/tpami.2022.3152247
- Oonishi K, Kunihiro N. Shor's Algorithm Using Efficient Approximate Quantum Fourier Transform. IEEE Transactions on Quantum Engineering. 2023;4:1-16. doi:10.1109/tqe.2023.3319044
- Aboy M, Marquez OW, McNames J, Hornero R, Trong T, Goldstein B. Adaptive Modeling and Spectral Estimation of Nonstationary Biomedical Signals Based on Kalman Filtering. IEEE Transactions on Biomedical Engineering. 2005;52(8):1485-1489. doi:10.1109/tbme.2005.851465
- Isik M, Vishwamith H, Inadagbo K, Dikmen IC. HPCNeuroNet: advancing neuromorphic audio signal processing with transformer-enhanced spiking neural networks. [preprint]. 2023. arXiv:2311.12449. doi:10.48550/arXiv.2311.12449.
- Leiber M, Marnissi Y, Barrau A, Badaoui ME. Differentiable Adaptive Short-Time Fourier Transform with Respect to the Window Length. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2023:1-5. doi:10.1109/icassp49357.2023.10095245
- Mastriani M. Quantum spectral analysis: frequency in time, with applications to signal and image processing. [preprint]. 2016. arXiv:1611.02302. doi:10.48550/arXiv.1611.02302.
- Jain PK, Raj Choudhary R, Singh MR. A Lightweight 1-D Convolution Neural Network Model for Multi-class Classification of Heart Sounds. 2022 International Conference on Emerging Techniques in Computational Intelligence (ICETCI). 2022:40-44. doi:10.1109/icetci55171.2022.9921376
- Wen P, Hu K, Yue W, Zhang S, Zhou W, Wang Z. Robust Audio Anti-Spoofing with Fusion-Reconstruction Learning on Multi-Order Spectrograms. INTERSPEECH 2023. 2023:271-275. doi:10.21437/interspeech.2023-563
- Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. 2017 IEEE International Conference on Computer Vision (ICCV). 2017:618-626. doi:10.1109/iccv.2017.74
- Baghel S, Prasanna SRM, Guha P. Overlapped speech detection using phase features. The Journal of the Acoustical Society of America. 2021;150(4):2770-2781. doi:10.1121/10.0006614
- Tuli S, Jha NK. EdgeTran: Device-Aware Co-Search of Transformers for Efficient Inference on Mobile Edge Platforms. IEEE Transactions on Mobile Computing. 2024;23(6):7012-7029. doi:10.1109/tmc.2023.3328287