Journal of Advanced Database Management & Systems

A Hybrid Algorithm for Duplicate Document Detection

  1. Ashish Kumar
  2. Arun Solanki

Abstract

Identification of duplicate document in a set of documents is a very big issue in information retrieval. In recent years, there are many researches going on and many methods have been proposed to detect and remove the duplicate documents but their relevance is still an issue. This paper proposed a hybrid algorithm based on word position by integrating n-gram searching technique. This experiment also uses inverted index to reduce the time complexity and space complexity and for fast searching. The result also shows a set of documents with their ranking thus helping the user not to waste time on getting the unwanted document. Cite this ArticleAshish Kumar, Arun Solanki. A Hybrid Algorithm for Duplicate Document Detection. Journal of Advanced Database Management & Systems. 2015; 2(2): 24–34p.

Keywords

Support