Search
13 articles for “Hadoop distributions”
-
Unveiling Distinctions: An In-depth Comparative Study of Cloudera, Microsoft Azure, and Amazon Web Services
Abstract: The purpose of this study is to compare and contrast the different on-premises and cloud-based Hadoop distributions. We will also analyze and contrast various cloud-based offerings and parts. We have settled on Cloudera as our Hadoop distribution of choice and Amazon and Azure as our cloud distribution services of choice. The advent of distributed cloud services has made it feasible to utilize and integrate a wide range of IT services …
Published in Journal of Advances in Shell Programming · Vol. 10, Issue 1, 2023 · pp. 27–34 Read article
-
Implementation of MRPrePost Parallel Algorithm based on Hadoop Platform for Large-Data Mining
Abstract: AbstractThe volume, velocity and variety of the data have increased several folds in the past few years. The conventional algorithms and techniques used to mine such huge data are found to be less efficient because these algorithms consider only the large threshold value due to which the number of candidates can be reduced, but this will lead mining association rules production to be inaccurate due to low utilization of data. …
Published in Recent Trends in Parallel Computing · Vol. 4, Issue 2, 2017 · pp. 10–20 Read article
-
Sentimental Analysis of Twitter Data
Abstract: Today, the world is rapidly heading towards the digital lifestyle. The data volume is getting larger and larger day by day. The solution for this problem can be resolved with the help of latest technology called as “Hadoop”. The data can be stored, processed, analyzed and later the results can be visualized. Here, in this paper, we are studying to fetch real time data from a social networking site, i.e. …
Published in Journal of Web Engineering & Technology · Vol. 6, Issue 1, 2019 · pp. 52–57 Read article
-
Distributed Data Mining Cloud Based System
Abstract: When traditional data mining systems are used for cloud computing, there are performance bottlenecks and scalability difficulties. In this study, it states that a data mining platform is directly based on the cloud computing. As compared with the traditional data mining systems, the platform is highly scalable, with massive data processing capabilities, service-oriented and low hardware costs. The platform can support the design and application of various distributed data mining …
Published in Journal of Operating Systems Development & Trends · Vol. 5, Issue 3, 2018 · pp. 10–14 Read article
-
Enhanced Incremental DB Scan Algorithm using Map-Reduce
Abstract: Distributed data mining is more efficient, scalable and its performance is better than the central data mining techniques. Incremental DBSCAN algorithm is better than the other method DBSCAN. Incremental DBSCAN can give better performance in distributed environment in terms of run time complexity using Hadoop platform. The proposed system uses spatial dataset i.e. dataset are distributed among different sites. Initially, clusters are generated at each site using DBSCAN algorithm. Then …
Published in Research & Reviews: A Journal of Embedded System & Applications · Vol. 3, Issue 2, 2015 · pp. 1–6 Read article
-
Effective Implementation of Apriori Algorithm to Develop a Suggestion System Based on Sales History Using Hadoop Environment
Abstract: AbstractThe need for comprehensive support system to analyze and predict the nature of the dynamic market based on the previous records is very vital in competent industries today. The data mining is a process of extracting implicit, previously unknown and potentially useful information from data. Mining is search for relationships and global patterns that exist in the large databases but that are hidden among vast amount of data. Since the …
Published in Journal of Advanced Database Management & Systems · Vol. 3, Issue 3, 2016 · pp. 8–16 Read article
-
Review Paper on Platform for Big Data Analytics as a Service with Stock Market Prediction and Analysis
Abstract: Petabyte-scale data is referred to as big data. The sum of the data in this quantity, which includes audio files, image files, and many more. Due to the unstructured nature of the data, it is challenging to evaluate the information that is gathered from it. As a result of social networking and cloud computing, data has significantly increased. It is therefore difficult to assess, process, and store. Big data processing …
Published in Journal of Advanced Database Management & Systems · Vol. 10, Issue 1, 2023 · pp. 1–5 Read article
-
Hadoop Yarn Big Data: Review Paper
Abstract: AbstractYARN, yet another resource negotiator is a large-scale, distributed operating system for big data applications and software rewrite that is capable of decoupling MapReduce's resource management and scheduling capabilities from the data processing component. In this study we have briefly described the evolution of YARN for big data analysis from the year 2017 to the recent times. Keywords: Apache HADOOP, HDFS, MapReduce, Apache Samza, Apache Flink
Published in Journal of Advanced Database Management & Systems · Vol. 7, Issue 3, 2020 · pp. 18–26 Read article
-
A Dynamic Text Compression Model for Big Data Applications Using Hadoop
Abstract: In today’s data-driven era, efficiently handling vast amounts of information has become increasingly important. Data compression plays a vital role in this regard — it is essentially a method of encoding information in such a way that significantly reduces the number of bits required to store or transmit a file. By shrinking data to its most compact form, compression techniques help save storage space, reduce bandwidth consumption, and improve the …
Published in International Journal of Algorithms Design and Analysis Review · Vol. 4, Issue 2, 2026 Read article
-
Hadoop and Map Reduce for Simplified Processing of Big Data
Abstract: Big data refers to large and complex data sets made up of a variety of structured and unstructured data, which are too big, too fast, or too hard to be managed by traditional techniques. Hadoop is an open source software platform for structuring big data on computer clusters built from commodity hardware and enabling it for data analysis. It is designed to scale up from a single server to thousands …
Published in Current Trends in Information Technology · Vol. 6, Issue 3, 2016 · pp. 1–6 Read article
-
Cloud Computing and Storage Issues
Abstract: This paper talks over security issues for cloud computing & presents a layered framework for secure clouds and then focus on layers, i.e., the storage layer & the data layer. Cloud computing is a set of IT facilities that are as long as to a customer over a network on adhoc basis and with the ability to scale up or down their service requirements. Generally cloud computing services are distributed …
Published in Recent Trends in Programming languages · Vol. 1, Issue 2, 2014 · pp. 20–26 Read article
-
Challenges in Parallel Computing for Big Data Analytics
Abstract: The integration of parallel computing into the realm of big data analytics promises accelerated processing speeds and enhanced scalability, but it is not without its formidable challenges. This study explores the multifaceted hurdles faced in the pursuit of efficient parallel processing for large-scale data analytics. The intricate task of distributing and partitioning massive datasets across multiple processing units demands adept strategies to ensure equitable workloads. Load balancing emerges as a …
Published in Recent Trends in Parallel Computing · Vol. 11, Issue 1, 2024 · pp. 1–6 Read article
-
Implementing FP-Growth Algorithm using Map Reduce for Mining Association Rules
Abstract: Abstract: In mining frequent itemsets, one of most important algorithms is FP-growth. FP-growth proposes an algorithm to compress information needed for mining frequent itemsets in FP-tree and recursively constructs FP-trees to find all frequent itemsets. Map Reduce is a distributed processing framework where the application is divided into many fragments of work, each of which may be executed on any node on a cluster. The main objective of this paper …
Published in Journal of Advanced Database Management & Systems · Vol. 6, Issue 2, 2019 · pp. 18–29 Read article