Search
28 articles for “audio processing”
-
Audio Summarization of Podcasts
Abstract: Podcasts have emerged as a significant medium for disseminating information, sharing stories, and providing entertainment. As their popularity continues to soar, the sheer volume of available content poses a challenge for listeners seeking to efficiently consume information. In this context, creating and deploying an audio summarizer for podcasts becomes highly significant. This research paper delves into the motivation for creating such a tool, emphasizing the increasing need for concise and …
Published in Journal of Software Engineering Tools & Technology Trends · Vol. 11, Issue 3, 2024 · pp. 18–26 Read article
-
Music Reactive led using Arduino
Abstract: In recent years, the integration of technology into interactive systems has garnered significant attention. Among the many applications, music-reactive LED systems have become popular, offering dynamic visualizations that respond to audio inputs. This paper explores the development and implementation of a music-reactive LED system using Arduino, focusing on real-time audio signal processing and LED control based on the frequency spectrum of the sound. By utilizing Fast Fourier Transform (FFT) algorithms …
Published in Journal of Microelectronics and Solid State Devices · Vol. 13, Issue 1, 2026 · pp. 27–33 Read article
-
Deep Learning Approach to Produce Artificial Speech (Text-To-Audio)
Abstract: This program utilizes key features of the .NET framework to facilitate smooth text-to-speech conversion and audio playback. Upon execution, users are prompted to input text via a graphical user interface (GUI), which the program converts into speech using the ‘SpeechSynthesizer’ class from the ‘System. Speech.Synthesis’ namespace. The audio that has been synthesized is handled and stored as a WAV file called ‘output.wav’ by utilizing the ‘FileStream’ class, allowing for future …
Published in Journal of Artificial Intelligence Research & Advances · Vol. 12, Issue 1, 2025 · pp. 28–33 Read article
-
Comparative Analysis Between Librosa and OpenSMILE
Abstract: This research work focuses on comparative study of Librosa, a python-based library, and openSMILE, a C++ toolkit, with python bindings used in audio speech analysis. Librosa is ideal for beginners due to its simple structure and flexibility with strong integration with machine learning frameworks like TensorFlow and PyTorch. On the other hand, OpenSMILE is ideal for speech-centric tasks like speech-emotion recognition or paralinguistic studies, offering a wide range of pre-defined …
Published in Journal of Open Source Developments · Vol. 12, Issue 3, 2025 · pp. 06–10 Read article
-
SpeakEasy: Python’s Desktop Companion for Effortless Interaction
Abstract: With their ability to use natural language processing to enable hands-free technological engagement, voice assistants have become an indispensable aspect of modern living. This paper describes how to create a voice assistant with the Python programming language by utilizing its powerful libraries and frameworks for task automation, speech recognition, and natural language processing. The system architecture makes use of several Python modules, including PyAudio, NLTK (Natural Language Toolkit), and Speech …
Published in Current Trends in Signal Processing · Vol. 14, Issue 1, 2024 · pp. 15–23 Read article
-
Identifying and Blocking of Non-productive Calls in Emergency Call System Using Machine Learning and IVRS Integration
Abstract: The Dial 100 emergency reaction gadget handles over 1,000,000 calls every day, with over 95% being unproductive, such as blank, machine-generated, and spoofed calls. These futile calls waste resources and put off responses to actual emergencies. This look presents a comprehensive solution integrating advanced name evaluation, system learning, and an interactive voice response system (IVRS) to filter out and prevent these calls. Our technique starts with studying incoming calls to …
Published in Journal of Communication Engineering & Systems · Vol. 14, Issue 3, 2024 · pp. 20–28 Read article
-
Voice Control Music System Using Raspberry Pi
Abstract: This paper presents a voice-controlled music system using Raspberry Pi, designed for hands-free music playback through voice commands. The system is implemented using Raspberry Pi 4B, a microphone, a speaker, and a 32GB SD card. The software is developed in Python using Thonny IDE, leveraging libraries such as SpeechRecognition, PyAudio, and gTTS for speech processing and audio playback. The system recognizes user commands like play, pause, next, and stop, converting …
Published in Journal of VLSI Design Tools and Technology · Vol. 15, Issue 2, 2025 · pp. 34–39 Read article
-
Advancements in AI-Driven Sound Spectrogram Analysis: From Deep Learning to Quantum and Neuromorphic Processing
Abstract: The rapid advancement of artificial intelligence (AI) has significantly reshaped the field of audio signal processing, with sound spectrogram analysis emerging as a central research focus. Spectrograms provide a rich time–frequency representation of audio signals, making them particularly suitable for data-driven learning approaches. This paper presents an in-depth and original review of modern AI-based techniques applied to spectrogram analysis, highlighting their growing impact across critical application areas such as healthcare …
Published in Journal of Multimedia Technology & Recent Advancements · Vol. 13, Issue 1, 2026 · pp. 01–06 Read article
-
AI Voice Detection Tool
Abstract: In today’s digital era, distinguishing between AI-generated and human voices is more important than ever. This project introduces an AI-based voice detection system designed to accurately identify synthetic voices, ensuring security and authenticity across various applications like cybersecurity, media verification, and fraud prevention.Our system works by analyzing incoming audio samples and comparing them against a diverse database of both AI-generated and real human voices. Using advanced machine learning and signal …
Published in Journal of Instrumentation Technology & Innovations · Vol. 16, Issue 1, 2026 · pp. 1–8 Read article
-
Robustness of Deepfake Detection Systems Against Adversarial Attacks
Abstract: This paper explores a deep learning system to detect deepfake videos, a common type of fake media. With the use of sophisticated methods such as recurrent neural networks (RNNs) and convolutional neural networks (CNNs), our system can reliably discern between authentic and altered videos. It analyzes both the images and the audio in videos to find signs of deepfake manipulation. We process video frames and audio, extract features with CNNs …
Published in Journal of Instrumentation Technology & Innovations · Vol. 14, Issue 2, 2024 Read article
-
Real-Time Multilanguage Platform with Encryption
Abstract: The proposed system, titled “Real-Time Multilanguage Platform With Encryption”, is designed to enable fast, reliable, and secure communication across various languages without relying on any third-party APIs. It utilizes a built-in audio-to-text conversion mechanism that processes spoken input using Python-based libraries like Speech Recognition or OpenAI’s Whisper, ensuring high accuracy in transcriptions. Once the audio is converted to text, the platform employs an offline translation engine to convert the transcribed …
Published in Journal of Communication Engineering & Systems · Vol. 15, Issue 2, 2025 · pp. 1–7 Read article
-
Avian Echoes: Convolutional Neural Network for Bird Vocalization Detection
Abstract: Bird species identification is a complex task within ornithology that demands advanced technological solutions. This research presents an approach leveraging Convolutional Neural Networks (CNNs) for bird species recognition based on identification of bird sound, each employing unique datasets and methodologies. The objective involves a two-stage identification process, beginning with the construction of an ideal dataset. The crucial step involves converting 1D audio waveforms to 2D spectrograms, enhancing CNNs' ability to …
Published in Journal of Aerospace Engineering & Technology · Vol. 14, Issue 2, 2024 · pp. 26–37 Read article
-
Innovative Applications of AI and Machine Learning in Voice-Driven Systems: Friday AI 2.0
Abstract: The development of artificial intelligence (AI) has enabled the development of increasingly sophisticated models that can understand and react to human questions. This study introduces Friday AI-2.0, a cutting-edge AI model that improves user engagement by using natural language processing (NLP) and voice synthesis technologies to produce text-based and audio-based responses. By increasing engagement through an intuitive and engaging interface, Friday AI-2.0 seeks to improve the user experience. By describing …
Published in Journal of Multimedia Technology & Recent Advancements · Vol. 12, Issue 3, 2025 · pp. 01–09 Read article
-
The Prospects of Multimedia in 2025: Emerging Patterns to Monitor
Abstract: The term multimedia describes the computer-aided combination of text, illustrations, graphics, audio, animation, still and moving images (videos), and any other medium that allows for the digital expression, storing, processing, and communication of any kind of information. The word of the decade is multimedia. It is a buzzword that has been used in a variety of contexts. It appears on the covers of movies, CDs, literature, newspapers, and handheld games. …
Published in Journal of Multimedia Technology & Recent Advancements · Vol. 11, Issue 3, 2024 · pp. 25–31 Read article
-
ChatGPT Based Voice Assistant for Blind People
Abstract: The proposed system for converting speech input into text format to facilitate interaction with ChatGPT is a sophisticated integration of hardware and cloud-based services. Utilizing state-of-the-art technologies, it facilitates seamless communication between users and the AI model. At the outset, the microphone serves as the input device, capturing audio signals from the user's speech. These signals are then amplified to ensure clarity and fidelity before being transmitted to the ESP32 …
Published in Journal of Operating Systems Development & Trends · Vol. 11, Issue 2, 2024 · pp. 23–31 Read article
-
Fault Diagnosis of Air Compressor (AC) System using Local Mean Decomposition (LMD) and Logistic Regression (LR) Machine Learning Classifier
Abstract: This article presents a detailed and systematic procedure for performing fault diagnosis in an air compressor (AC) system by analyzing the audio signals generated during its operation. The analysis covers both normal (healthy) conditions and seven distinct types of faults, including bearing failure, flywheel malfunction, inlet valve leakage, outlet valve leakage, non-return valve failure, piston ring defect, and rider belt issues. To acquire the acoustic signals, the researchers utilized a …
Published in Journal of Polymer & Composites · Vol. 14, Issue 1, 2026 · pp. 416–427 Read article
-
Multimodal Generative AI for Vehicular Applications at Edge
Abstract: This paper explores how the rapid advancements in vehicular technology, including autonomous driving and intelligent transportation systems, have driven the need for real-time data processing and decision-making. Multimodal generative AI, when deployed at the edge, offers a powerful solution for vehicular applications by leveraging diverse data sources such as video, audio, sensor inputs, and environmental data. This paper explores the integration of multimodal generative AI with edge computing in vehicular …
Published in Journal of Control & Instrumentation · Vol. 16, Issue 1, 2025 · pp. 27–34 Read article
-
An Adaptive and Privacy-Aware Federated Learning Framework for Efficient and Secure Model Training Across Heterogeneous Datasets
Abstract: The problem of efficiency and privacy regarding heterogeneous data in modern distributed machine learning systems is a vital point that should be taken into account. The absence of IID data distribution, client heterogeneity, and privacy invasion during the aggregation model are the bane of conventional federated learning (FL) approaches to learning like FedAvg and FedProx. The paper proposes that the adaptive and privacy-aware FL framework (AFL-P) can be used to …
Published in Journal of Mobile Computing, Communications & Mobile Networks · Vol. 13, Issue 1, 2026 · pp. 16–25 Read article
-
Comparative Analysis of MCNN and RCNN for Speech Emotion Recognition Using Gender Information
Abstract: Speech emotion recognition is a speech processing task and a computer-based approach designed to identify and classify the emotions conveyed in audio signals. The aim of this system is to evaluate a speaker's emotional state, such as happiness, anger, sadness, or frustration, by analyzing their speech patterns, which include prosodic features like pitch, frequency, and rhythm. Speech emotion recognition is used in various real-life scenarios that include Customer Service, Healthcare, …
Published in Journal of Communication Engineering & Systems · Vol. 15, Issue 1, 2025 · pp. 1–10 Read article
-
Survey Paper on Multilingual Live Call Translation Using Deep Learning
Abstract: This research work surveys cutting-edge language translation technologies, including multi-lingual, real-time translation, voice recognition, speech-to-text conversion, and transcription in the hearing process. The study explores the complex mechanisms behind voice call language translation, focusing on sophisticated machine learning models integrated with cloud-based or local applications to facilitate seamless communication across language barriers. Furthermore, conducting research in live communication analyzes the complexity of text and voice techniques to deliver translated content …
Published in Journal of Image Processing & Pattern Recognition Progress · Vol. 11, Issue 2, 2024 · pp. 13–21 Read article