Search
8 articles for “GPUs”
-
Parallelization of Metaheuristics for the Optimization of Permuted Perceptron Problem
Abstract: Parallel computing has found a mainstream backing in graphics processing units (GPUs).These resources have enormous processing capacity, are energy efficient, and are broadly available, unlike grids. Since the advent of CUDA (Compute Unified Device Architecture) developed by NVIDIA that permits GPU programming in C, C++ language, GPUs are now being used in a variety of fields, including scientific computing. This thesis focuses on the optimization of the solution of a …
Published in Journal of Operating Systems Development & Trends Read article
-
Energy-efficient Image Classification on Edge Devices: Implementation and Evaluation
Abstract: Image classification is a computer vision problem where an algorithm determines a class or label for a given image. Various real-time applications like object recognition, medical diagnosis, person recognition, etc. Image classification property on edge devices is useful for autonomous vehicles, surveillance, and healthcare and internet of things deployments. The advancement of deep learning based methods and graphics processing units (GPU) devices allows efficient processing locally. The study utilizes a …
Published in Journal of Image Processing & Pattern Recognition Progress · Vol. 11, Issue 3, 2024 · pp. 10–18 Read article
-
Automating Compiler Optimization: A Machine Learning Approach
Abstract: This study reports on an ML-based approach to compiler optimization, complementing traditional optimization methods that rely strongly on hand-tuned settings. Compiler optimization plays a key role in performance-speedup and energy optimization of complex contemporary software systems. However, the traditional approach to optimizer settings involves laborious, error-prone, and scale-insensitive human-in-the-loop intervention, especially in the complex and high-demand environments in which today's computing application thrives. By integrating RL and GA, we can …
Published in Journal of Artificial Intelligence Research & Advances · Vol. 12, Issue 1, 2025 · pp. 12–16 Read article
-
TensorFlow: Architecture, Applications, and Future Challenges
Abstract: TensorFlow, an open-source machine learning platform created by Google, has revolutionized how artificial intelligence (AI) systems are built and implemented. Designed to support scalable and flexible model training across CPUs, GPUs, and TPUs, TensorFlow enables researchers and developers to construct advanced deep learning models with efficiency and precision. This study provides an in-depth examination of TensorFlow's architecture, including its use of dataflow graphs and tensor-based computation. We explore its adaptability …
Published in Journal of Open Source Developments · Vol. 12, Issue 2, 2025 · pp. 41–50 Read article
-
Design and Optimization of Domain-Specific Languages for High-Performance Computing Applications
Abstract: The accelerating demand for computational power in scientific, engineering, and data-intensive domains has driven high-performance computing (HPC) systems toward unprecedented levels of parallelism and architectural complexity. Contemporary HPC platforms integrate multicore CPUs, many-core GPUs, accelerators, and deep memory hierarchies, creating significant challenges for software development and performance optimization. Traditional general-purpose programming languages and parallel programming frameworks provide low-level control over hardware resources but require extensive manual tuning, resulting in poor …
Published in Recent Trends in Parallel Computing · Vol. 13, Issue 1, 2026 · pp. 17–22 Read article
-
A Review Study on CPU-Optimized Parameter-Efficient Fine-Tuning for Large Language Models to Increase Accuracy Using LoRA
Abstract: The fast proliferation of large language models (LLMs) has increased the need to optimize the process of fine-tuning, but the existing workflows that require a graphics processing unit (GPU) are still expensive, intensive, and unavailable to most researchers. This paper is driven by the desire to have a more cost-efficient and democratized version by examining a CPU-efficient implementation of parameter-efficient fine-tuning (PEFT) based on low-rank adaptation (LoRA). The major purpose …
Published in Recent Trends in Parallel Computing · Vol. 13, Issue 1, 2026 · pp. 32–38 Read article
-
The Evolution of Digital Integrated Circuits: A Journey Through Technological Advancements
Abstract: The evolution of digital integrated circuits (ICs) is a remarkable journey that has transformed the worldof technology over the past several decades. This abstract provides an overview of the comprehensivearticle that explores this evolution. The journey begins with the birth of digital integrated circuits in thelate 1950s and early 1960s when the first ICs containing a small number of transistors emerged. Theseearly ICs laid the groundwork for a technological revolution …
Published in Journal of Microcontroller Engineering and Applications · Vol. 10, Issue 2, 2023 · pp. 15–18 Read article
-
A Five-Layer Architectural Framework for Sustainable and Scalable AI Systems
Abstract: Artificial Intelligence (AI) is not only about algorithms. AI works like a full “stack” of layers, from electricity to real-world user applications. In this paper, we explain a simple and student-friendly Five- Layer Architecture of AI: (1) Energy, (2) Chips, (3) Infrastructure, (4) Models, and (5) Applications. Each layer supports the next layer, like a cake with multiple layers. If any layer is weak, AI systems become slow, costly, or …
Published in Journal of Artificial Intelligence Research & Advances · Vol. 13, Issue 3, 2025 Read article