Journal of Multimedia Technology & Recent Advancements Review Article

Advancements in AI-Driven Sound Spectrogram Analysis: From Deep Learning to Quantum and Neuromorphic Processing

  1. Sanjeev Sharma Department of Multimedia, BBK DAV College For Women, Amritsar

Abstract

The rapid advancement of artificial intelligence (AI) has significantly reshaped the field of audio signal processing, with sound spectrogram analysis emerging as a central research focus. Spectrograms provide a rich time–frequency representation of audio signals, making them particularly suitable for data-driven learning approaches. This paper presents an in-depth and original review of modern AI-based techniques applied to spectrogram analysis, highlighting their growing impact across critical application areas such as healthcare diagnostics, security systems, environmental surveillance, and intelligent multimedia processing. We examine a range of state-of-the-art methodologies, including convolutional neural networks (CNNs), transformer-based models, and adaptive spectral estimation strategies, emphasizing how each approach leverages spectro-temporal patterns for improved feature extraction and classification. Through multiple case studies – covering deepfake audio detection, disease diagnosis from biomedical sounds, and acoustic scene classification – we demonstrate that AI-driven models consistently outperform conventional signal processing methods in terms of accuracy, generalization, and noise robustness. A key contribution of this work is the discussion of explainable artificial intelligence (XAI) techniques, which enhance model transparency and trust by revealing how spectral features influence decision-making. Experimental evaluations show that transformer architectures equipped with spectrogram-aware attention mechanisms achieve up to 92.4% accuracy on standard benchmark datasets, exceeding CNN-based models by 6.8%. Finally, the paper identifies current technical challenges and outlines future research directions, including multimodal data fusion, efficient edge deployment, and the potential role of quantum-enhanced spectral analysis in next-generation audio intelligence systems.

Keywords

References (30)

  1. Anagha R, Arya A, Narayan VH, Abhishek S, Anjali T. Audio Deepfake Detection Using Deep Learning. 2023 12th International Conference on System Modeling & Advancement in Research Trends (SMART). 2023:176-181. doi:10.1109/smart59791.2023.10428163
  2. Balamurugan A, Teo SG, Yang J, Peng Z, Xulei Y, Zeng Z. ResHNet: Spectrograms Based Efficient Heart Sounds Classification Using Stacked Residual Networks. 2019 IEEE EMBS International Conference on Biomedical & Health Informatics (BHI). 2019:1-4. doi:10.1109/bhi.2019.8834578
  3. Nogueira AFR, Oliveira HS, Machado JJM, Tavares JMRS. Transformers for Urban Sound Classification—A Comprehensive Performance Evaluation. Sensors. 2022;22(22):8874. doi:10.3390/s22228874
  4. Jones CC, Gannon ZE, Blunt SD, Allen CT, Martone AF. An Adaptive Spectrogram Estimator to Enhance Signal Characterization. 2022 IEEE Radar Conference (RadarConf22). 2022:1-6. doi:10.1109/radarconf2248738.2022.9764186
  5. Tsui BMW, Xu J, Rittenbach A, Chen S, El-Sharkaway AM, Edelstein WA, et al. High performance SPECT system for simultaneous SPECT-MR imaging of small animals. 2011 IEEE Nuclear Science Symposium Conference Record. 2011:3178-3182. doi:10.1109/nssmic.2011.6153652
  6. Tian B, Pang Y, Huzaifa M, Wang S, Adve S. Towards Energy-Efficiency by Navigating the Trilemma of Energy, Latency, and Accuracy. 2024 IEEE International Symposium on Mixed and Augmented Reality (ISMAR). 2024:913-922. doi:10.1109/ismar62088.2024.00107
  7. Xia Y, Zhao Z. Cross-modal Background Suppression for Audio-Visual Event Localization. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022:19957-19966. doi:10.1109/cvpr52688.2022.01936
  8. Abdul ZK, Al-Talabani AK. Mel Frequency Cepstral Coefficient and its Applications: A Review. IEEE Access. 2022;10:122136-122158. doi:10.1109/access.2022.3223444
  9. Cerezuela-Escudero E, Jimenez-Fernandez A, Paz-Vicente R, Dominguez-Morales M, Linares-Barranco A, Jimenez-Moreno G. Musical notes classification with neuromorphic auditory system using FPGA and a convolutional spiking network. 2015 International Joint Conference on Neural Networks (IJCNN). 2015:1-7. doi:10.1109/ijcnn.2015.7280619
  10. Li F, Zhang Z, Wang L, Liu W. Heart sound classification based on improved mel-frequency spectral coefficients and deep residual learning. Frontiers in Physiology. 2022;13. doi:10.3389/fphys.2022.1084420
  11. Kim G, Han DK, Ko H. SpecMix: a mixed sample data augmentation method for training with time-frequency domain features. [preprint]. 2021. arXiv:2108.03020. doi:10.48550/arXiv.2108.03020.
  12. Foresti GL, Regazzoni CS. Multisensor data fusion for autonomous vehicle navigation in risky environments. IEEE Transactions on Vehicular Technology. 2002;51(5):1165-1185. doi:10.1109/tvt.2002.800629
  13. Wang H, Zou Y, Wang W. SpecAugment++: a hidden space data augmentation method for acoustic scene classification. [preprint]. 2021. arXiv:2103.16858. doi:10.48550/arXiv.2103.16858.
  14. Han J, Matuszewski M, Sikorski O, Sung H, Cho H. Randmasking Augment: A Simple and Randomized Data Augmentation For Acoustic Scene Classification. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2023:1-5. doi:10.1109/icassp49357.2023.10095001
  15. Thuillier E, Gamper H, Tashev IJ. Spatial Audio Feature Discovery with Convolutional Neural Networks. 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2018:6797-6801. doi:10.1109/icassp.2018.8462315
  16. Schluter J, Gutenbrunner G. EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use. 2022 30th European Signal Processing Conference (EUSIPCO). 2022:205-208. doi:10.23919/eusipco55093.2022.9909910
  17. Nguyen TTM, Nguyen DD, Luong CM. Vietnamese Speaker Verification With Mel-Scale Filter Bank Energies and Deep Learning. IEEE Access. 2024;12:150114-150122. doi:10.1109/access.2024.3479092
  18. Wang J, Li J, Tan X. Spectral-Spatial Symmetrical Aggregation Cross-Linking Multi-Modal Data Fusion Network. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2022:5098-5102. doi:10.1109/icassp43922.2022.9747570
  19. Barahona S, de Benito-Gorrón D, Toledano DT, Ramos D. Enhancing Conformer-Based Sound Event Detection Using Frequency Dynamic Convolutions and BEATs Audio Embeddings. IEEE/ACM Transactions on Audio, Speech, and Language Processing. 2024;32:3896-3907. doi:10.1109/taslp.2024.3444490
  20. Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, et al. A Survey on Vision Transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2023;45(1):87-110. doi:10.1109/tpami.2022.3152247
  21. Oonishi K, Kunihiro N. Shor's Algorithm Using Efficient Approximate Quantum Fourier Transform. IEEE Transactions on Quantum Engineering. 2023;4:1-16. doi:10.1109/tqe.2023.3319044
  22. Aboy M, Marquez OW, McNames J, Hornero R, Trong T, Goldstein B. Adaptive Modeling and Spectral Estimation of Nonstationary Biomedical Signals Based on Kalman Filtering. IEEE Transactions on Biomedical Engineering. 2005;52(8):1485-1489. doi:10.1109/tbme.2005.851465
  23. Isik M, Vishwamith H, Inadagbo K, Dikmen IC. HPCNeuroNet: advancing neuromorphic audio signal processing with transformer-enhanced spiking neural networks. [preprint]. 2023. arXiv:2311.12449. doi:10.48550/arXiv.2311.12449.
  24. Leiber M, Marnissi Y, Barrau A, Badaoui ME. Differentiable Adaptive Short-Time Fourier Transform with Respect to the Window Length. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2023:1-5. doi:10.1109/icassp49357.2023.10095245
  25. Mastriani M. Quantum spectral analysis: frequency in time, with applications to signal and image processing. [preprint]. 2016. arXiv:1611.02302. doi:10.48550/arXiv.1611.02302.
  26. Jain PK, Raj Choudhary R, Singh MR. A Lightweight 1-D Convolution Neural Network Model for Multi-class Classification of Heart Sounds. 2022 International Conference on Emerging Techniques in Computational Intelligence (ICETCI). 2022:40-44. doi:10.1109/icetci55171.2022.9921376
  27. Wen P, Hu K, Yue W, Zhang S, Zhou W, Wang Z. Robust Audio Anti-Spoofing with Fusion-Reconstruction Learning on Multi-Order Spectrograms. INTERSPEECH 2023. 2023:271-275. doi:10.21437/interspeech.2023-563
  28. Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. 2017 IEEE International Conference on Computer Vision (ICCV). 2017:618-626. doi:10.1109/iccv.2017.74
  29. Baghel S, Prasanna SRM, Guha P. Overlapped speech detection using phase features. The Journal of the Acoustical Society of America. 2021;150(4):2770-2781. doi:10.1121/10.0006614
  30. Tuli S, Jha NK. EdgeTran: Device-Aware Co-Search of Transformers for Efficient Inference on Mobile Edge Platforms. IEEE Transactions on Mobile Computing. 2024;23(6):7012-7029. doi:10.1109/tmc.2023.3328287
Support