Search
9 articles for “Apache Hadoop”
-
Review Paper on Platform for Big Data Analytics as a Service with Stock Market Prediction and Analysis
Abstract: Petabyte-scale data is referred to as big data. The sum of the data in this quantity, which includes audio files, image files, and many more. Due to the unstructured nature of the data, it is challenging to evaluate the information that is gathered from it. As a result of social networking and cloud computing, data has significantly increased. It is therefore difficult to assess, process, and store. Big data processing …
Published in Journal of Advanced Database Management & Systems · Vol. 10, Issue 1, 2023 · pp. 1–5 Read article
-
Decoding Big Data: A Practical Comparison Between Hadoop and Spark
Abstract: This paper conducts a comprehensive comparison of Apache Hadoop and Apache Spark, two essential frameworks in the big data era. The rapid expansion of data possesses challenges in terms of volume, variety, and velocity, which necessitate advanced processing solutions. Hadoop, utilizing its MapReduce paradigm, provides scalable and fault-tolerant storage, whereas Spark, built upon Hadoop, introduces in-memory processing to increase speed and flexibility. This study includes a detailed examination of their …
Published in Recent Trends in Parallel Computing · Vol. 11, Issue 3, 2024 · pp. 15–23 Read article
-
Big Data: A Survey Paper on Big Data Technologies
Abstract: AbstractThere are many Big Data technologies that have been making an impact on the new technology stacks for handling Big Data, but Apache Hadoop is one technology that has been the darling of Big Data talk. Hadoop is an open-source platform for storage and processing of diverse data types that enables data-driven enterprises to rapidly derive the complete value from all their data. A plethora of Big Data Analytics technologies …
Published in Journal of Computer Technology & Applications · Vol. 11, Issue 3, 2020 · pp. 1–7 Read article
-
Effects of Cluster Computing on Big Data Analysis and Network Topology
Abstract: The rapid expansion of big data has posed substantial difficulties for conventional computing systems. As a result, cluster computing has grown to be a potent method for effective large data processing. Cluster computing involves multiple interconnected nodes functioning as a unified system, pooling together their processing, storage, and memory resources. These nodes are typically connected through high-speed networks such as ethernet or InfiniBand, facilitating efficient data sharing and communication among …
Published in International Journal of Algorithms Design and Analysis Review · Vol. 1, Issue 1, 2023 · pp. 31–39 Read article
-
Hadoop Yarn Big Data: Review Paper
Abstract: AbstractYARN, yet another resource negotiator is a large-scale, distributed operating system for big data applications and software rewrite that is capable of decoupling MapReduce's resource management and scheduling capabilities from the data processing component. In this study we have briefly described the evolution of YARN for big data analysis from the year 2017 to the recent times. Keywords: Apache HADOOP, HDFS, MapReduce, Apache Samza, Apache Flink
Published in Journal of Advanced Database Management & Systems · Vol. 7, Issue 3, 2020 · pp. 18–26 Read article
-
Big Data Tools: A Survey
Abstract: AbstractNowadays, a large volume of data is generated in the form of text, voice, video, images, and sound. It is a very challenging job to handle and to get processed these different types of data. It is a very laborious process to analyze big data by using traditional data processing applications. Due to huge scattered file systems, a big data analysis is a difficult task. So, to analyze big data, …
Published in E-Commerce for Future & Trends · Vol. 7, Issue 3, 2020 · pp. 23–34 Read article
-
Challenges in Parallel Computing for Big Data Analytics
Abstract: The integration of parallel computing into the realm of big data analytics promises accelerated processing speeds and enhanced scalability, but it is not without its formidable challenges. This study explores the multifaceted hurdles faced in the pursuit of efficient parallel processing for large-scale data analytics. The intricate task of distributing and partitioning massive datasets across multiple processing units demands adept strategies to ensure equitable workloads. Load balancing emerges as a …
Published in Recent Trends in Parallel Computing · Vol. 11, Issue 1, 2024 · pp. 1–6 Read article
-
Review Paper on Analyzing the Data using Bigdata and Hadoop
Abstract: AbstractIn this world of information, the term ‘Big data’ has emerged with new opportunities and challenges to deal with massive amount of data. Big data has earned a place of great importance and is becoming the choice for new researches. This paper presents an overview on Big data, technical challenges and Hadoop. An overview on technologies used by big data application to handle the massive data are Hadoop, Map Reduce, …
Published in E-Commerce for Future & Trends · Vol. 7, Issue 3, 2020 · pp. 1–9 Read article
-
Developing Intelligent Business Solutions Using Big Data
Abstract: Different types of industries and organizations are generating a large volume of data (known as big data). Analysis of such data has the potential to provide meaningful business insights and can be used to make various business decisions. However, the extraction of meaningful information from big data is not an easy task as it comes with different types of challenges. Production of actionable information and profitable decisions from big data …
Published in Journal of Software Engineering Tools & Technology Trends · Vol. 7, Issue 1, 2020 · pp. 12–17 Read article