1 publication
-
Published Subscription
A Hybrid Algorithm for Duplicate Document DetectionBy Ashish Kumar, Arun Solanki
Abstract: Identification of duplicate document in a set of documents is a very big issue in information retrieval. In recent years, there are many researches going on and many methods have been proposed to detect and remove the duplicate documents but their relevance is still an issue. This paper proposed a hybrid algorithm based on word position by integrating n-gram searching technique. This experiment also uses inverted index to reduce the …
Published in Journal of Advanced Database Management & Systems · Vol. 2, Issue 2, 2015 · pp. 24–34 Read article →