Recent Trends in Parallel Computing Review Article

Decoding Big Data: A Practical Comparison Between Hadoop and Spark

  1. Saunved Palve Department of Computer Engineering, Shri Vile Parle Kelavani Mandal's Narsee Monjee Institute of Management Studies (SVKM'S NMIMS), Navi Mumbai
  2. Ammar Abdulhussain Department of Computer Engineering, Shri Vile Parle Kelavani Mandal's Narsee Monjee Institute of Management Studies (SVKM'S NMIMS), Navi Mumbai
  3. Om Sawant Department of Computer Engineering, Shri Vile Parle Kelavani Mandal's Narsee Monjee Institute of Management Studies (SVKM'S NMIMS), Navi Mumbai
  4. Ayush Yadav Department of Computer Engineering, Shri Vile Parle Kelavani Mandal's Narsee Monjee Institute of Management Studies (SVKM'S NMIMS), Navi Mumbai
  5. Aditya Kasar School of Technology, Shri Vile Parle Kelavani Mandal's Narsee Monjee Institute of Management Studies (SVKM'S NMIMS), Navi Mumbai

Abstract

This paper conducts a comprehensive comparison of Apache Hadoop and Apache Spark, two essential frameworks in the big data era. The rapid expansion of data possesses challenges in terms of volume, variety, and velocity, which necessitate advanced processing solutions. Hadoop, utilizing its MapReduce paradigm, provides scalable and fault-tolerant storage, whereas Spark, built upon Hadoop, introduces in-memory processing to increase speed and flexibility. This study includes a detailed examination of their features, strengths, and limitations, offering insights through experiments in statistical analysis, machine learning, and database operations. The work contributes valuable perspectives for both practitioners and researchers, enabling them to make informed decisions in the ever-changing landscape of big data analytics. Our experiments reveal that Apache Spark achieves a notable average speedup of 41.57% over Apache Hadoop in statistical and machine learning applications, underscoring its superiority in big data analytics. However, it is important to note that Hadoop excels in database management systems, demonstrating superior performance in particular scenarios.

Keywords

References (18)

  1. Anjum B. MapReduce–The scalable distributed data processing solution. In: Topics in Parallel and Distributed Computing: Enhancing the Undergraduate Curriculum: Performance, Concurrency, and Programming on Modern Platforms. 2018. p. 173–90.
  2. Naga Malleswari TYJ, Vadivu G. MapReduce: A Technical Review. Indian Journal of Science and Technology. 2016;9(1). doi:10.17485/ijst/2016/v9i1/78964
  3. Feller E, Ramakrishnan L, Morin C. On the performance and energy efficiency of Hadoop deployment models. 2013 IEEE International Conference on Big Data. 2013:131-136. doi:10.1109/bigdata.2013.6691564
  4. Hong Z, Xiao-Ming W, Jie C, Yan-Hong M, Yi-Rong G, Min W. An optimized model for MapReduce based on Hadoop. TELKOMNIKA. 2016;14:1552–8.
  5. Kala Karun A, Chitharanjan K. A review on hadoop — HDFS infrastructure extensions. 2013 IEEE CONFERENCE ON INFORMATION AND COMMUNICATION TECHNOLOGIES. 2013:132-137. doi:10.1109/cict.2013.6558077
  6. Salloum S, Dautov R, Chen X, Peng PX, Huang JZ. Big data analytics on Apache Spark. International Journal of Data Science and Analytics. 2016;1(3-4):145-164. doi:10.1007/s41060-016-0027-9
  7. Wang K, Khan MMH. Performance Prediction for Apache Spark Platform. 2015 IEEE 17th International Conference on High Performance Computing and Communications, 2015 IEEE 7th International Symposium on Cyberspace Safety and Security, and 2015 IEEE 12th International Conference on Embedded Software and Systems. 2015:166-173. doi:10.1109/hpcc-css-icess.2015.246
  8. Shaikh E, Mohiuddin I, Alufaisan Y, Nahvi I. Apache Spark: A Big Data Processing Engine. 2019 2nd IEEE Middle East and North Africa COMMunications Conference (MENACOMM). 2019:1-6. doi:10.1109/menacomm46666.2019.8988541
  9. García-Gil D, Ramírez-Gallego S, García S, Herrera F. A comparison on scalability for batch big data processing on Apache Spark and Apache Flink. Big Data Analytics. 2017;2(1). doi:10.1186/s41044-016-0020-2
  10. Joshi A, Luo Y, John LK. Applying Statistical Sampling for Fast and Efficient Simulation of Commercial Workloads. IEEE Transactions on Computers. 2007;56(11):1520-1533. doi:10.1109/tc.2007.70748
  11. Aravinth SS, Begam AH, Shanmugapriyaa S, Sowmya S, Arun E. An efficient HADOOP frameworks SQOOP and Ambari for big data processing. Int J Innov Res Sci Technol. 2015;1: 252–5.
  12. Benlachmi Y, El Yazidi A, Hasnaoui ML. A comparative analysis of Hadoop and Spark frameworks using word count algorithm. Int J Adv Comput Sci Appl. 2021;12:778–88.
  13. Benbrahim H, Hachimi H, Amine A. Comparison between Hadoop and Spark. Proceedings of the International Conference on Industrial Engineering and Operations Management; 2019 Mar 5-7; Bangkok, Thailand. IEOM Society International; 2019. p. 690–701.
  14. Meng X, Bradley J, Yavuz B, Sparks E, Venkataraman S, Liu D, et al. Mllib: Machine learning in Apache Spark. J Mach Learn Res. 2016;17:1–7.
  15. Shi J, Qiu Y, Minhas UF, Jiao L, Wang C, Reinwald B, et al. Clash of the titans. Proceedings of the VLDB Endowment. 2015;8(13):2110-2121. doi:10.14778/2831360.2831365
  16. Tekdogan T, Cakmak A. Benchmarking Apache Spark and Hadoop MapReduce on Big Data Classification. 2021 5th International Conference on Cloud and Big Data Computing (ICCBDC). 2021:15-20. doi:10.1145/3481646.3481649
  17. Assoicate professor CS IT, Jain Univrsity, Bangalore., Pavan DKU, Nachappa DMN, Professor and Academic Head, School of CS and IT, Jain University., Srinivasu DSVN, Professor CSE, Narasaraopeta Engineering College (Autonomous), Narasaraopet, A.P. Sqoop usage in Hadoop Distributed File System and Observations to Handle Common Errors. International Journal of Recent Technology and Engineering (IJRTE). 2020;9(4):452-454. doi:10.35940/ijrte.d4980.119420
  18. Verma A, Mansuri AH, Jain N. Big data management processing with Hadoop MapReduce and spark technology: A comparison. 2016 Symposium on Colossal Data Analysis and Networking (CDAN). 2016:1-4. doi:10.1109/cdan.2016.7570891
Support