Journal of Mobile Computing, Communications & Mobile Networks Review Article

A Supervised Learning Approach for Toxic Comment Detection on Social Media Platforms

  1. M. Prasad Department of Computer Science Engineering, Shri Vishnu Engineering College for Women (A), Bhimavaram
  2. Rajarao PBV Department of Computer Science Engineering, Shri Vishnu Engineering College for Women (A), Bhimavaram
  3. P. Kiran Sree Department of Computer Science Engineering, Shri Vishnu Engineering College for Women (A), Bhimavaram
  4. P. Gowthami Department of Computer Science Engineering, Shri Vishnu Engineering College for Women (A), Bhimavaram
  5. K. Ajita Lakshmi Department of Electronics and Communication Engineering, Shri Vishnu Engineering College for Women(A), Bhimavaram
  6. B. Subrahmanyam Department of MCA, BVC Institute of Technology and Science(A), Batlapalem

Abstract

Nowadays everyone uses social media platforms like X (formerly Twitter), Instagram, Facebook, etc. for various purposes. With the help of this, we share our opinions, ideas, and feelings. Generally, the datasets obtained from the internet are constructive; however, there is a significant proportion of toxic ones. The datasets are filtered to remove noise, and noise is removed in post-processing. The study initiates with the upload and preprocessing of a toxic comment dataset meticulously cleaning text by eliminating stop words and special symbols to lay out a standardized corpus. The resulting application of count vectorizer captures word occurrences constructing a feature matrix for algorithmic training supervised algorithms, including support vector machine, logistic regression, naive Bayes, random forest, decision tree, and k-nearest neighbors. These are systematically implemented and each algorithm undergoes rigorous assessment with accuracy measurements computed to check its proficiency in segregating toxic from non-toxic comments. The examination finishes in an accuracy graph visually contrasting the performance of the various supervised algorithms. This visual representation helps in identifying the most effective model for online toxic comment classification. The multiheaded model consists of toxicity, severe-toxic, obscene threat insults, and toxicity prediction based on confusion metrics. The practical implications of this study lie in outfitting a robust tool for online platforms to automatically detect and manage toxic comments adding to a safer and more constructive digital environment.

Keywords

References (12)

  1. Badjatiya P, Gupta S, Gupta M, Varma V. Deep learning for hate speech detection in tweets. In: Proceedings of the 26th International Conference on World Wide Web Companion, Perth, Australia, April 3–7, 2017. pp. 759–760.
  2. Prusa J, Khoshgoftaar TM, Dittman DJ, Napolitano A. Using random undersampling to alleviate class imbalance on tweet sentiment data. In: Proceedings of the IEEE International Conference on Information Reuse and Integration, August 13–15, 2015. pp. 197–202.
  3. Waseem Z, Davidson T, Warmsley D, Weber I. A typology of abusive language detection subtasks for understanding abuse. In: Proceedings of the the First Workshop on Abusive Language Online. Vancouver, British Columbia, Canada: Association for Computational Lingistics; 2017. pp. 78–84.
  4. Rustam F, Khalid M, Aslam W, Rupapara V, Mehmood A, Choi GS. A performance comparison of supervised machine learning models for COVID-19 tweets sentiment analysis. PLoS One. 2021; 16 (2): e0245909.
  5. Fatima EB, Omar B, Abdelmajid EM, Rustam F, Mehmood A, Choi GS. Minimizing the overlapping degree to improve class imbalanced learning under sparse feature selection: application to fraud detection. IEEE Access. 2021; 9: 28101–28110.
  6. Rustam F, Ashraf I, Mehmood A, Ullah S, Choi GS. Tweets classification on the base of sentiments for US airline companies. Entropy. 2019; 21 (11): 1078.
  7. Anandarajan M, Hill C, Nolan T. Practical Text Analytics: Maximizing the Value of Text Data. Advances in Analytics and Data Science, vol. 2. Cham, Switzerland: Springer; 2019.
  8. Raja Rao PBV, Prasad M, Kiran Sree P, Venkata Ramana C, Satyanarayana Murty PT. Enhancing the MANET AODV Forecast of a Broken Link with LBP. Smart Innovation, Systems and Technologies. 2023:51-69. doi:10.1007/978-981-99-4717-1_6
  9. Maddula P, Srikanth P, Sree PK, Rao PBVR, Murty PTS. COVID-19 prediction with Chest X-Ray images using CNN. 2023 International Conference on Intelligent and Innovative Technologies in Computing, Electrical and Electronics (IITCEE). 2023:568-572. doi:10.1109/iitcee57236.2023.10090951
  10. Prasad M, Lakshmi KA, Pbv RR, Prasanthi BV, Sree PK, Babu VR, et al. A CNN and TF Techniques Development for Efficient Identification of Floral Recognition. 2024 IEEE International Conference on Computing, Power and Communication Technologies (IC2PCT). 2024:327-332. doi:10.1109/ic2pct60090.2024.10486528
  11. Satyanarayana Murty PT, Prasad M, Raja Rao PBV, Kiran Sree P, Ramesh Babu G, Phaneendra Varma C. A Hybrid Intelligent Cryptography Algorithm for Distributed Big Data Storage in Cloud Computing Security. Lecture Notes in Computer Science. 2023:637-648. doi:10.1007/978-3-031-36402-0_59
  12. Sree PK, Chintalapati PV, N SUD, Prasad M, Babu GR, Raja Rao PBV. Waste Management Detection Using Deep Learning. 2023 3rd International Conference on Computing and Information Technology (ICCIT). 2023:50-54. doi:10.1109/iccit58132.2023.10273898
Support