Journal of Advanced Database Management & Systems

Identification of Thyroid Disease Severity using Fuzzy C-Means

  1. G. Jaya Suma
  2. M. Nymisha
  3. Sk. Teenaaz
  4. S. Neelima
  5. K. Gouri Shankar

Abstract

Clustering is a primary data description method in data mining which group’s most similar data. Data clustering is one of the important problems in the fields of data mining, bio-informatics and pattern recognition. Various algorithms are used to solve this problem. This paper presents the performance analysis of k-means clustering algorithm and compares it with fuzzy C- Means (FCM) algorithm on Thyroid disease data set. We also find the accuracy of both the algorithms on Thyroid data set. Fuzzy clustering is useful to mine multi-dimensional and complex data sets, where the members have fuzzy or partial relations. Among the various techniques developed so far, fuzzy-C-Means (FCM) algorithm is the most popular one, where a partial membership is assigned to a piece of data with each of the pre-defined cluster centers. Moreover, in FCM, the cluster centers are virtual, that is, they are chosen at random and thus might be out of the data set. With in some iteration the cluster centers and membership values of the data points are updated. The FCM employs fuzzy portioning such that a point can belong to more than one group with different membership grades between 0 and 1. Using this membership we are going to determine the severity of thyroid disease in a person.Keywords: K-means, fuzzy C-means clustering, Euclidean distance, Manhattan distance, sigmoid membership, t-norms
Support