International Journal of Computer Science Languages Original Research
Comparative Study of BERT Variants for Sentiment Analysis with Error Analysis
Abstract
The use of media is going up fast in India, and this has led to the rise of Hinglish. Hinglish is an informal blend of Hindi and English that people commonly use in everyday conversations, especially across social media platforms such as Twitter, Facebook, and WhatsApp. People use Hinglish to talk to each other in a way that is not very formal. Hinglish blends English vocabulary with informal usage, often ignoring standard grammatical rules, which makes it challenging for computers to accurately interpret and process it. It is especially hard for computers to figure out how people are feeling when they use Hinglish in the media. Hinglish poses a significant challenge for natural language processing, particularly when it comes to accurately interpreting people’s emotions and opinions. The problem with Hinglish is that it often has words from languages mixed together, and the spelling and sentence structure can be weird. This makes it hard for regular language models to understand. Even though models like BERT are really good at understanding text in languages, they need a lot of computer power to work. This means they are not good for situations where we need to get answers, and we do not have a lot of computer power. So, we looked at how three smaller models work: DistilBERT, Multilingual Representations for Indian Languages (MuRIL), and XLM-RoBERTa. We used the Kaggle Hinglish Sentiment Dataset to test these models. When we look at how these models work and where they make mistakes, the research helps us understand how they can handle the complexities of Hinglish language. This is important because Hinglish is a mix of Hindi and English. The research shows that these models can work with Hinglish while still being good at classifying things and not using much computer power. Studying is helpful because it adds to what we know about working with languages that do not have a lot of resources. It also helps us make systems that can figure out how people feel about things. We can use these systems in real life. The research on Hinglish models is useful for sentiment analysis systems.
Keywords
References (19)
- Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. In: Burstein J, Doran C, Solorio T, editors. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Vol. 1, Long and Short Papers; 2019 Jun; Minneapolis, MN, USA. Stroudsburg (PA): Association for Computational Linguistics; 2019. p. 4171–4186. doi:10.18653/v1/N19-1423.
- Çetinoğlu Ö, Schulz S, Vu NT. Challenges of Computational Processing of Code-Switching. Proceedings of the Second Workshop on Computational Approaches to Code Switching. 2016:1-11. doi:10.18653/v1/w16-5801
- Patil A, Patwardhan V, Phaltankar A, Takawane G, Joshi R. Comparative Study of Pre-Trained BERT Models for Code-Mixed Hindi-English Data. 2023 IEEE 8th International Conference for Convergence in Technology (I2CT). 2023:1-7. doi:10.1109/i2ct57861.2023.10126273
- Hashmi E, Yayilgan SY, Shaikh S. Augmenting sentiment prediction capabilities for code-mixed tweets with multilingual transformers. Social Network Analysis and Mining. 2024;14(1). doi:10.1007/s13278-024-01245-6
- Astuti LW, Sari Y, Suprapto. Code-Mixed Sentiment Analysis using Transformer for Twitter Social Media Data. International Journal of Advanced Computer Science and Applications. 2023;14(10). doi:10.14569/ijacsa.2023.0141053
- Sampath KK, Supriya M. Transformer Based Sentiment Analysis on Code Mixed Data. Procedia Computer Science. 2024;233:682-691. doi:10.1016/j.procs.2024.03.257
- Tewari P, Gumber M, Tyagi S, Seekhwal P. Hinglish Text Analysis: Challenges and Opportunities in Multilingual Natural Language Processing. SSRN Electronic Journal. 2025. doi:10.2139/ssrn.5193470
- Jadon AS, Parmar M, Agrawal R. Hinglish Sentiment Analysis: Deep Learning Models for Nuanced Sentiment Classification in Multilingual Digital Communication. 2024 2nd International Conference on Device Intelligence, Computing and Communication Technologies (DICCT). 2024:318-323. doi:10.1109/dicct61038.2024.10533057
- Singh SK, Sharma A, Sahil, Singh D, Pandit S, Saghir U. Sentiment Analysis of English-Hindi Code-Mixed Text Using mBERT Model. 2025 3rd International Conference on Inventive Computing and Informatics (ICICI). 2025:552-556. doi:10.1109/icici65870.2025.11069692
- Mohana Priya K.T, Shrinithi G, Nithish P, Panesh A C. COMPARATIVE ANALYSIS OF TRANSFORMER MODELS FOR SENTIMENT CLASSIFICATION IN CODE- MIXED INDIC LANGUAGES. International Journal of Engineering Research and Sustainable Technologies (IJERST). 2025;3(1):1-9. doi:10.63458/ijerst.v3i1.101
- Almalki SS. Sentiment Analysis and Emotion Detection Using Transformer Models in Multilingual Social Media Data. International Journal of Advanced Computer Science and Applications. 2025;16(3). doi:10.14569/ijacsa.2025.0160332
- Aliyu Y, Sarlan A, Danyaro KU, Sani abd Rahman A, Muazu AA, Abubakar MY. Deep learning techniques for sentiment analysis in code-switched Hausa-English tweets. International Journal of Information Management Data Insights. 2025;5(1):100330. doi:10.1016/j.jjimei.2025.100330
- Ramesh G, Doddapaneni S, Bheemaraj A, Jobanputra M, AK R, Sharma A, et al. Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages. Transactions of the Association for Computational Linguistics. 2022;10:145-162. doi:10.1162/tacl_a_00452
- Nazir MK, Faisal CN, Habib MA, Ahmad H. Leveraging Multilingual Transformer for Multiclass Sentiment Analysis in Code-Mixed Data of Low-Resource Languages. IEEE Access. 2025;13:7538-7554. doi:10.1109/access.2025.3527710
- Mamta, Ekbal A. Transformer based multilingual joint learning framework for code-mixed and english sentiment analysis. Journal of Intelligent Information Systems. 2023;62(1):231-253. doi:10.1007/s10844-023-00808-x
- Veeramani H, Thapa S, Naseem U. MLInitiative@WILDRE7: hybrid approaches with large language models for enhanced sentiment analysis in code-switched and code-mixed texts. In: Jha GN, Sobha L, Bali K, Ojha AK, editors. Proceedings of the 7th Workshop on Indian Language Data: Resources and Evaluation; 2024 May; Torino, Italia. Paris: ELRA and ICCL; 2024. p. 66–72.
- Thakur V, Sahu R, Omer S. Current State of Hinglish Text Sentiment Analysis. SSRN Electronic Journal. 2020. doi:10.2139/ssrn.3614442
- Yuan LS, Ming LT. Sentiment Prediction Using Multilingual Bidirectional Encoder Representations And Cross-Lingual Language Model Robustly Optimized Bert Approach From Transformers on Code-mixed Text. 2025 IEEE International Conference on Computation, Big-Data and Engineering (ICCBE). 2025:901-905. doi:10.1109/iccbe65177.2025.11256168
- Kumar A, Susan S. Supervised Sentiment Analysis of Movie Reviews with SHAP-Based Interpretability Analysis. Lecture Notes in Networks and Systems. 2025:381-388. doi:10.1007/978-981-96-3361-6_28