
Nabarun Halder
- Research Assistant, CCDS
Research Interests
Machine Learning, Deep Learning, Natural Language Processing, Health Informatics, Human-Computer Interaction
Jahanggir Hossain Setu, Nabarun Halder, Sankar Sikder, Ashraful Islam, Md Zahangir Alam
2024 International Joint Conference on Neural Networks (IJCNN)
IEEE, pp. 1-7
Arabic Natural Language Processing (NLP) presents unique challenges due to the complexity of the language, including variations in dialects, rich morphology, and context-dependent semantics. These linguistic intricacies make accurate classification a challenging task. This study presents the development and evaluation of an Arabic Bidirectional Encoder Representations from Transformers (AraBERT) model trained on a publicly available dataset named UltimateArabic dataset, which comprises 10 diverse classes. The aim is to perform multiclass classification and to address the imbalance class problem, using AraVec, of Arabic text data in the field of NLP. The model achieved an overall accuracy of 98% on augmented data, showcasing its proficiency in categorizing a wide range of Arabic texts including medical, technology, art, culture, diversity, society, religion, sports, politics, and economy. The AraBERT model exhibits strong classification performance compared to non-augmented data across the majority of categories, with an encouraging raise on macro-average precision of 4%, recall of 4%, and F1-score of 4%. This study underscores the potential of AraBERT along with AraVec for diverse applications in imbalanced Arabic text classification tasks and highlights the need for continued research to address the complexities of multiclass classification in Arabic NLP.

Machine Learning, Deep Learning, Natural Language Processing, Health Informatics, Human-Computer Interaction

Assistant Professor
Department of Computer Science and Engineering
Independent University, Bangladesh (IUB)Human-Computer Interaction, AI for Social Good, AI for Public Health, AI for Impact