outlier detection
1 article · search the full text for this term
-
Improving Dataset Integrity Through Automated Data Cleaning Techniques
Abstract: High-quality data is a fundamental requirement in data science for producing trustworthy analytical insights and effective machine learning models. Problems, including incomplete records, inconsistent entries, duplicate observations, and anomalous values, can severely reduce the accuracy and robustness of predictive systems. As modern datasets continue to expand in both volume and structural complexity, relying on manual data cleaning methods become time-consuming and error-prone, highlighting the growing importance of automated data preprocessing …
Published in International Journal of Data Structure Studies · Vol. 4, Issue 1, 2026 · pp. 40–45 Read article