
Shadman Rohan
- Research Assistant, CCDS
Research Interests
Machine Learning, Natural Language Processing
Shadman Rohan, Mahmud Elahi Akhter, Ibraheem Muhammad Moosa, Nabeel Mohammed, Amin Ahsan Ali, Akmmahbubur Rahman
Proceedings of the Second Workshop on Bangla Language Processing (BLP-2025)
In: Proceedings of the Second Workshop on Bangla Language Processing (BLP-2025)
Association for Computational Linguistics, pp. 347-356
We study sentence-level generative data augmentation for Bangla semantic classification across four public datasets and three pretrained model families (BanglaBERT, XLM-Indic, mBERT).We evaluate two widely used, reproducible techniquesparaphrasing (mT5-based) and round-trip backtranslation (BnEnBn)and analyze their impact under realistic class imbalance.Overall, augmentation often helps, but gains are tightly coupled to label quality: paraphrasing typically outperforms backtranslation and yields the most consistent improvements for the monolingual model, whereas multilingual encoders benefit less and can be more sensitive to noisy minority-class expansions.A key empirical observation is that the neutral class appears to be a major source of annotation noise, which degrades decision boundaries and can cap the benefits of augmentation even when positive/negative classes are clean and polarized.We provide practical guidance for Bangla sentiment pipelines: (i) use simple sentence-level augmentation to rebalance classes when labels are reliable; (ii) allocate additional curation and higher interannotator agreement targets to the neutral class.Our results indicate when augmentation helps and suggest that data qualitynot model choice alonecan become the limiting factor.

Machine Learning, Natural Language Processing

Professor
Department of Computer Science and Engineering
Independent University, Bangladesh
Artificial Intelligence, Machine Learning