Search
16 articles for “Audio Learning Model”
-
Automatic Baby Cry Detector with sleep music player (ABCD)
Abstract: In today’s world, our lives have become moredependent on technology. One problem which caught our eyesis when parents have to leave their wards (aged between 3months to 2 years) alone due to some essential tasks for a shortduration, the baby goes unmonitored. Normally, parents dothis by leaving their child asleep. In many cases, the childwakes up and starts to cry. In absence of loved ones, babiesneed immediate calmness relief. A …
Published in Journal of Mechatronics and Automation · Vol. 10, Issue 1, 2023 · pp. 21–29 Read article
-
Lip Reading: Transforming Speech to Text
Abstract: Lip reading, the ability to interpret spoken language by observing lip movements, is a valuable skill that can aid in various applications, particularly in enhancing speech recognition systems. This project explores the implementation of a deep learning-based lip-reading model to improve the accuracy and robustness of speech recognition in challenging environments, such as noisy or audio-limited settings. The proposed lip-reading system leverages Convolutional Neural Networks (CNNs) and Recurrent Neural Networks …
Published in Current Trends in Signal Processing · Vol. 14, Issue 1, 2024 · pp. 23–33 Read article
-
Classification Of Music Genre Using Machine Learning
Abstract: Machine Learning is a way that helps the systems to automatically learn from experience and improve the performance and predict the outcome more accurately without any requirement of being explicitly programmed. Music is one of the most significant and influential part of people’s life. Also, music is known as a universal language as it has the power to unite people from different places and cultures. This helps in the recognition …
Published in Journal of Instrumentation Technology & Innovations · Vol. 12, Issue 1, 2022 · pp. 11–16 Read article
-
Acoustic Sensing for City Flow: Quasi-Supervised Recognition of Sirens and Traffic for Urban Mobility Intelligence
Abstract: This paper frames environmental audio as a mobility telemetry source, extending a benchmark urban-sound corpus with transportation-critical classes—ambulance, firetruck, police, and traffic—and training spectrogram-based models under a quasi-supervised regime to support real-time city operations; leveraging 10-fold protocols, class-weighted objectives, and audiospecific augmentations (time stretch, pitch shift, SpecAugment, PatchAugment), the system benchmarks multiple CNN backbones combined with self-supervised learning paradigms enable the extraction of rich, discriminative acoustic representations, achieving strong multi-class …
Published in Trends in Electrical Engineering · Vol. 16, Issue 1, 2025 · pp. 42–50 Read article
-
Classifying Abnormalities in Heartbeat Sound
Abstract: Heartbeat sounds play a major role in the detection of various diseases such as heart disease, hyperthyroidism, and high blood pressure in their early stages. In the proposed method, various abnormal and healthy heartbeat audio signals are given as input and the features are extracted using MFCC (mel-frequency cepstral coefficients). Then, a deep learning approach is applied in which the MFCC audio signals are sent to the CNN (convolutional neural …
Published in Research & Reviews: A Journal of Embedded System & Applications · Vol. 12, Issue 1, 2024 · pp. 24–31 Read article
-
Survey Paper on Multilingual Live Call Translation Using Deep Learning
Abstract: This research work surveys cutting-edge language translation technologies, including multi-lingual, real-time translation, voice recognition, speech-to-text conversion, and transcription in the hearing process. The study explores the complex mechanisms behind voice call language translation, focusing on sophisticated machine learning models integrated with cloud-based or local applications to facilitate seamless communication across language barriers. Furthermore, conducting research in live communication analyzes the complexity of text and voice techniques to deliver translated content …
Published in Journal of Image Processing & Pattern Recognition Progress · Vol. 11, Issue 2, 2024 · pp. 13–21 Read article
-
ChatGPT Based Voice Assistant for Blind People
Abstract: The proposed system for converting speech input into text format to facilitate interaction with ChatGPT is a sophisticated integration of hardware and cloud-based services. Utilizing state-of-the-art technologies, it facilitates seamless communication between users and the AI model. At the outset, the microphone serves as the input device, capturing audio signals from the user's speech. These signals are then amplified to ensure clarity and fidelity before being transmitted to the ESP32 …
Published in Journal of Operating Systems Development & Trends · Vol. 11, Issue 2, 2024 · pp. 23–31 Read article
-
An Adaptive and Privacy-Aware Federated Learning Framework for Efficient and Secure Model Training Across Heterogeneous Datasets
Abstract: The problem of efficiency and privacy regarding heterogeneous data in modern distributed machine learning systems is a vital point that should be taken into account. The absence of IID data distribution, client heterogeneity, and privacy invasion during the aggregation model are the bane of conventional federated learning (FL) approaches to learning like FedAvg and FedProx. The paper proposes that the adaptive and privacy-aware FL framework (AFL-P) can be used to …
Published in Journal of Mobile Computing, Communications & Mobile Networks · Vol. 13, Issue 1, 2026 · pp. 16–25 Read article
-
AI Voice Detection Tool
Abstract: In today’s digital era, distinguishing between AI-generated and human voices is more important than ever. This project introduces an AI-based voice detection system designed to accurately identify synthetic voices, ensuring security and authenticity across various applications like cybersecurity, media verification, and fraud prevention.Our system works by analyzing incoming audio samples and comparing them against a diverse database of both AI-generated and real human voices. Using advanced machine learning and signal …
Published in Journal of Instrumentation Technology & Innovations · Vol. 16, Issue 1, 2026 · pp. 1–8 Read article
-
Advancements in AI-Driven Sound Spectrogram Analysis: From Deep Learning to Quantum and Neuromorphic Processing
Abstract: The rapid advancement of artificial intelligence (AI) has significantly reshaped the field of audio signal processing, with sound spectrogram analysis emerging as a central research focus. Spectrograms provide a rich time–frequency representation of audio signals, making them particularly suitable for data-driven learning approaches. This paper presents an in-depth and original review of modern AI-based techniques applied to spectrogram analysis, highlighting their growing impact across critical application areas such as healthcare …
Published in Journal of Multimedia Technology & Recent Advancements · Vol. 13, Issue 1, 2026 · pp. 01–06 Read article
-
AI and Machine Learning Approaches for Estimating Depression Severity: Techniques, Trends, and Applications
Abstract: Depression is a very common mental health disorder that results in a disorder of a person’s behavior, emotions, and cognitive abilities. Depression can be caused by environmental factors or hereditary factors. The person suffering from depression might have symptoms of suicidal thoughts, altering food patterns as well as sleeping issues. Depression is a global issue that has impacted millions of people globally having more effect on women worldwide. The complexity …
Published in Journal of Electronic Design Technology · Vol. 15, Issue 3, 2024 · pp. 29–38 Read article
-
Innovative Applications of AI and Machine Learning in Voice-Driven Systems: Friday AI 2.0
Abstract: The development of artificial intelligence (AI) has enabled the development of increasingly sophisticated models that can understand and react to human questions. This study introduces Friday AI-2.0, a cutting-edge AI model that improves user engagement by using natural language processing (NLP) and voice synthesis technologies to produce text-based and audio-based responses. By increasing engagement through an intuitive and engaging interface, Friday AI-2.0 seeks to improve the user experience. By describing …
Published in Journal of Multimedia Technology & Recent Advancements · Vol. 12, Issue 3, 2025 · pp. 01–09 Read article
-
Automatic Car Controller Based on Sign Board using Deep Learning and IOT
Abstract: The rapid growth of intelligent transportation systems has increased the demand for safer and more efficient driving solutions. Conventional vehicles often rely heavily on human intervention, which can lead to accidents due to negligence, fatigue, or poor visibility of traffic signs. This project proposes an automated car control system that utilizes deep learning and Internet of Things (IoT) technologies to recognize traffic signboards and respond accordingly. The primary objective is …
Published in Journal of Control & Instrumentation · Vol. 17, Issue 2, 2026 · pp. 16–24 Read article
-
Face Detection and Recognition Using MTCNN and FaceNet
Abstract: Face detection and face recognition are major tasks in the field of computer vision with several real-world applications and many products being developed in the same field. This study gives a detailed implementation of the product that is developed for accurate detection and recognition of faces along with audio output of the face detected. This development would act as a base for a few future products that can be developed …
Published in Journal of Artificial Intelligence Research & Advances · Vol. 11, Issue 2, 2024 · pp. 132–140 Read article
-
Digital Audio Watermarking: An Overview
Abstract: AbstractA real watermarking procedure is required for copyright security and confirmation of protected innovation. This work recommends an advanced watermarking procedure which makes utilization of concurrent recurrence veiling to conceal the watermark data. The calculation is grounded on Psychoacoustic auditory model and Spread spectrum hypothesis. It creates a watermark sign utilizing spread range hypothesis and implants it into the sign by measuring the covering edge utilizing modified Psychoacoustic auditory model. …
Published in Journal of Communication Engineering & Systems · Vol. 5, Issue 1, 2015 · pp. 1–9 Read article
-
A Review Of Deep Learning Applications For Speech Processing Improvement
Abstract: Improve the quality of the spoken word is a common goal for many audio and speech signal processing applications. A noisy voice signal's quality and understandability may be improved via speech augmentation. Speech augmentation is critical in a wide range of fields, including hearing aids, ASR, and mobile communication. DNN-based architectures for speech recognition and augmentation have shown to be quite effective in recent years, according to a new study. …
Published in Journal of Telecommunication, Switching Systems and Networks · Vol. 9, Issue 2, 2022 · pp. 14–19 Read article