Gender Abusive Language Detection in Bengali Using Machine Learning Algorithms
Published in Proceedings of the 2nd International Conference on Big Data, IoT and Machine Learning
Part of the book series: Lecture Notes in Networks and Systems ((LNNS, volume 867))
Included in the following conference series: International Conference on Big Data, IoT and Machine Learning
The issue of gender-based abuse is a prevalent concern in today’s society. As technology and social media platforms have become increasingly ubiquitous, these platforms have also become a breeding ground for abusive language and harassment, particularly toward women. In this study, we utilized machine learning techniques—logistic regression, decision tree, random forest, K-nearest neighbors, support vector machine, and Naïve Bayes—to classify abusive text based on gender. The considered dataset in this research comprised comments and posts from various social media platforms, which were preprocessed before being subjected to classification. Experimental analysis revealed that support vector machine demonstrated superior performance in terms of precision, recall, accuracy, sensitivity, and specificity indicating its potential effectiveness in identifying and filtering out gender-based abuse from social media platforms. The findings of this study suggest that machine learning techniques can play a critical role in combating gender-based abuse and harassment online.
Gender Abusive Bengali Text Classification using Enhanced CNN-LSTM Model with Bangla BERT Base Preprocessing
Accepted and presented on 27th International Conference on Computer and Information Technology.
Abstract:
With the rapid proliferation of social media and online communication platforms, the presence of gender-based abusive language has emerged as a significant issue in the digital sphere. Addressing this challenge necessitates the development of effective Natural Language Processing (NLP) techniques capable of accurately detecting and categorizing abusive content to safeguard users from harmful interactions. This study introduces a novel approach for Gender-Based Abusive Bengali Text Classification, which plays a crucial role in fostering safer online environments. The proposed methodology employs an enhanced CNN-LSTM model in conjunction with Bangla BERT Base for preprocessing, resulting in an impressive accuracy of 97.94%, outperforming prior models and demonstrating its effectiveness in identifying gender-targeted abusive language within the Bengali-speaking community. The primary contribution of this work lies in the development of a robust and precise classifier for gender-based abusive Bengali text, which holds significant potential for mitigating harmful online behavior. Consequently, this research not only advances the domain of NLP but also supports the creation of a safer and more inclusive digital space for Bengali language users.
Analyzing Audience Engagement in Esports: Sentiment and LLM-Based Topic Insights from Live Chats in South Asia
Accepted and presented on 27th International Conference on Computer and Information Technology.
Abstract:
This study investigates the dynamics of audience engagement in esports through sentiment analysis of live chat data and topic discovery, focusing on the popular game PUBG Mobile across Bangladesh, India, and Pakistan. A dataset encompassing nearly 15 million live chat messages and video metadata was utilized, employing a pre-trained RoBERTa model for sentiment classification, categorizing user sentiments into positive, negative, and neutral. The analysis revealed a significant increase in positive sentiment among Bangladeshi viewers, suggesting a shift towards a more favorable perception of esports. Additionally, the Gemma 7b-it large language model was applied to identify key discussion topics within the live chat, uncovering themes related to gameplay strategies and community interactions. The findings indicate a strong correlation between view counts and audience engagement, highlighting opportunities for advertisers to connect with dedicated esports fans. Despite limitations such as the focus on official YouTube channels and the resource constraints of sentiment analysis, this research offers valuable insights into real-time audience engagement in esports, paving the way for future studies to explore broader contexts and multimodal data integration.
Classifying Fake News with a Large-Scale Dataset Using Machine Learning Approaches
Manuscript under review, submitted to "2025 IEEE International Conference on Quantum Photonics, Artificial Intelligence, and Networking".