Recent Trends in Parallel Computing Review Article
A Comparative Review of Parallel Computing Techniques for High-Performance Computing
Abstract
High-Performance Computing (HPC) has become an essential technology for solving computationally intensive problems in scientific computing, engineering, artificial intelligence, weather forecasting, computational biology, financial modelling, and large-scale data analytics. The increasing complexity and heterogeneity of modern HPC systems have created a strong requirement for efficient parallel computing techniques and portable programming models. Several programming approaches, including Message Passing Interface (MPI), Open Multi-Processing (OpenMP), Compute Unified Device Architecture (CUDA), OpenACC, SYCL, Kokkos, RAJA, and hybrid MPI-based approaches, have been developed to exploit parallel hardware resources. This review examines and compares the key parallel computing techniques widely used in modern high-performance computing (HPC) systems. Recent research published between 2020 and 2026 is examined with emphasis on execution performance, speedup, scalability, parallel efficiency, communication overhead, programming complexity, portability, memory utilization, energy efficiency, and hardware support. The reviewed studies indicate that MPI remains highly effective for distributed memory and large-scale cluster computing, whereas OpenMP provides a relatively simple approach for shared-memory multicore systems. CUDA can provide highly optimized GPU performance but is strongly associated with NVIDIA architecture. Directive-based OpenACC simplifies accelerator programming, while SYCL provides a promising single-source and multi-vendor programming approach. Kokkos and RAJA offer performance-portable abstractions for developing applications across diverse architectures. Hybrid MPI+X approaches are particularly important for modern heterogeneous supercomputers. The review identifies a research gap in the lack of a unified framework that evaluates these programming models using common performance, portability, programmability, scalability, communication, and energy-related criteria. The study concludes that no single parallel programming technique is universally optimal; instead, the appropriate model depends on hardware architecture, application characteristics, workload, communication requirements, and portability objectives.
Keywords
References (17)
- Czarnul P, Proficz J, Drypczewski K. Survey of methodologies, approaches, and challenges in parallel programming using high-performance computing systems. Sci Program. 2020;2020(1):4176794.
- Lai J, Yu H, Tian Z, Li H. Hybrid MPI and CUDA parallelization for CFD applications on multi-GPU HPC clusters. Sci Program. 2020;2020(1):8862123.
- Kim JY, Kang JS, Joh M. GPU acceleration of MPAS microphysics WSM6 using OpenACC directives: performance and verification. Comput Geosci. 2021 Jan;146:104627.
- Maris P, Yang C, Oryspayev D, Cook B. Accelerating an iterative eigensolver for nuclear structure configuration interaction calculations on GPUs using OpenACC. J Comput Sci. 2022 Mar;59:101554.
- Shin W, Yoo KH, Baek N. Large-scale data computing performance comparisons on SYCL heterogeneous parallel processing layer implementations. Appl Sci (Basel). 2020 Mar;10(5):1656.
- Kuncham GK, Vaidya R, Barve M. Performance study of GPU applications using SYCL and CUDA on Tesla V100 GPU. In: 2021 IEEE High Performance Extreme Computing Conference (HPEC); 2021 Sep 20. p. 1–7.
- Reguly IZ. Comparative evaluation of bandwidth-bound applications on the Intel Xeon CPU Max Series. In: Proceedings of the SC'23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis; 2023 Nov 12. p. 1236–1244.
- Trott C, Berger-Vergiat L, Poliakoff D, Rajamanickam S, Lebrun-Grandie D, Madsen J, et al. The Kokkos ecosystem: comprehensive performance portability for high performance computing. Comput Sci Eng. 2021;23(5):10–18.
- Artigues V, Kormann K, Rampp M, Reuter K. Evaluation of performance portability frameworks for the implementation of a particle-in-cell code. Concurr Comput Pract Exp. 2020;32(11):e5640.
- Johnston B, Vetter JS, Milthorpe J. Evaluating the performance and portability of contemporary SYCL implementations. In: 2020 IEEE/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC); 2020 Nov 13. p. 45–56.
- Marowka A. On the performance portability of OpenACC, OpenMP, Kokkos and RAJA. In: International Conference on High Performance Computing in Asia-Pacific Region; 2022 Jan 7. p. 103–114.
- Breyer M, Van Craen A, Pflüger D. A comparison of SYCL, OpenCL, CUDA, and OpenMP for massively parallel support vector machine classification on multi-vendor hardware. In: Proceedings of the 10th International Workshop on OpenCL; 2022 May 10. p. 1–12.
- García Sánchez C, El Faqir El Rhazoui Y. Exploring the performance and portability of the k-means algorithm on SYCL across CPU and GPU architectures. J Supercomput. 2023;79:18480–18506. doi:10.1007/s11227-023-05373-2.
- Crisci L, Carpentieri L, Cosenza B, Accordi G, Gadioli D, Vitali E, et al. Enabling performance portability on the LiGen drug discovery pipeline. Future Gener Comput Syst. 2024 Sep;158:44–59.
- Costanzo M, Rucci E, García-Sánchez C, Naiouf M, Prieto-Matías M. Analyzing the performance portability of SYCL across CPUs, GPUs, and hybrid systems with SW sequence alignment. Future Gener Comput Syst. 2025 Sep;170:107838.
- Marowka A. Evaluating SYCL as a unified programming model for heterogeneous systems. Parallel Comput. 2026 Apr 11;103195.
- Krishnasamy E, Throtter J, Cai X, Pleiter D, Kos L, Saavedra L, et al. Performance and programmability of MPI+X integration with CUDA, HIP, SYCL, OpenACC, and OpenMP offloading for supercomputing: a case study on dense matrix-vector multiplication. In: Proceedings of the Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops; 2026 Jan 26. p. 457–468.