Conferences

Arabic Offense Text Detection in Social Networks using Collaborative Machine Learning

2024

Abstract

In an era of abundant social media use and digital communication, the classification of Arabic social media content and user comments into relevant categories poses a significant challenge. The content in social network platforms often encompasses a wide range of topics, sentiments, and societal concerns, making it essential to develop effective methods for discerning and categorizing these expressions. In this research, a new dataset for Arabic offensive text is built along with the methodology proposed for classifying Arabic content into five categories collected from different social networking sites. Leveraging a diverse array of text representation techniques, including TF-IDF, Bag of Words (BoW), custom embeddings, and pre-trained embeddings from models such as Arabert and Mbert, a comprehensive evaluation is conducted based, shedding light on the optimal approach for tackling this complex problem. The findings reveal that the TF-IDF and BoW methods stand out with an impressive 89% accuracy, emphasizing their robustness in addressing the nuances of Arabic tweet classification across these sensitive topics. This study offers valuable insights into the challenges inherent in categorizing Arabic social media content, emphasizing the significance of the TF-IDF approach as an effective tool for this task.