Journal of Web Engineering & Technology

Pattern Discovery for Text Mining

  1. Ravindra Changala
  2. D Rajeswara Rao
  3. Vidyullatha P
  4. Annapurna Gummadi

Abstract

Text mining refers to the process of extracting interesting, non-trivial information and knowledge from unstructured text. The challenging issue is to find accurate cognition or characteristics in text documents to help users to find what they want. Many term-based methods solved this but later suffered from the problems of polysemy and synonymy where polysemy means a word has multiple meanings and synonymy is multiple words having the same meaning. Later phrase-based approaches served well, as phrases may carry more “semantics” like information. The performance of it decreases due to phrases having inferior statistical properties to terms, low frequency of occurrence, and large numbers of redundant and noisy phrases among them. To overcome this pattern mining-based approaches have been proposed, which include the concept of closed sequential patterns, and pruned non-closed patterns. These pattern mining-based approaches worked with effectiveness. The contradiction is people think pattern-based approaches are significant alternative, but less improvements is made for the effectiveness compared with term-based methods.Keywords: Text mining, closed sequential patterns, pruned non-closed patterns, taxonomy discovery model, term based, phrase based
Support