Obfuscation-aware Cyberbullying Detection using BERTweet with Conditional Leet-speak Normalization
DOI:
https://doi.org/10.21928/uhdjst.v10n2y2026.pp134-152Keywords:
Cyberbullying Detection, Leet-Speak Normalization, Transformer Encoder, Obfuscation Robustness, Text Classification, Social Media AnalysisAbstract
Cyberbullying through digital platforms is a major concern on social media. Pretrained transformer encoders have performed well on detection tasks involving clean social-media text; however, little research has examined their robustness when characters are substituted with visually similar digits or symbols. This study proposes an obfuscation-aware framework consisting of a transformer encoder pretrained on Twitter data and a leet-speak normalization pipeline. The rule-based detector checks whether each input contains leet-speak; if leet-speak is detected, a character-level normalizer processes the input before encoding; otherwise, the input is passed to the encoder unchanged. The framework was evaluated on two publicly available English cyberbullying datasets, each containing 15,592 instances. Six encoders from three pretraining domains were evaluated on clean and synthetically obfuscated data using three random seeds. The results show that all six encoders degraded under the evaluated synthetic leet-speak transformations, with drops of up to 48% points relative to their clean-text baselines. The proposed pipeline improves leet_test macro-F1 by 18.13% points on the first dataset and by 13.92% points on the second, without reducing performance on clean text. A heuristic lower-bound analysis was used to assess whether the gains remained above the baseline across the tested random seeds.
References
A. G. Philipo, D. S. Sarwatt, J. Ding, M. Daneshmand and H. Ning. “Cyberbullying detection: Exploring datasets, technologies, and approaches on social media platforms”. ACM Computing Surveys, vol. 58, pp. 1-35, 2025.
C. Raj, A. Agarwal, G. Bharathy, B. Narayan and M. Prasad. “Cyberbullying detection: Hybrid models based on machine learning and natural language processing techniques”. Electronics, vol. 10, no. 22, p. 2810, 2021.
M. Umer, E. A. Alabdulqader, A. A. Alarfaj, L. Cascone and M. Nappi. “Cyberbullying detection using PCA extracted GLOVE features and RoBERTaNet transformer learning model”. IEEE Transactions on Computational Social Systems, vol. 12, no. 5, pp. 3881-3890, 2025.
D. Q. Nguyen, T. Vu and A. T. Nguyen. “BERTweet: A pre-trained language model for English tweets”. In: EMNLP 2020 - Conference on Empirical Methods in Natural Language Processing, Proceedings of Systems Demonstrations, pp. 9-14, 2020.
I. V. De Mendizabal, X. Vidriales, V. Basto-Fernandes, E. Ezpeleta, J. R. Méndez and U. Zurutuza. “Deobfuscating leetspeak with deep learning to improve spam filtering”. International Journal of Interactive Multimedia and Artificial Intelligence, vol. 8, no. 4, pp. 46-55, 2023.
T. Gröndahl, L. Pajola, M. Juuti, M. Conti and N. Asokan. “All you Need is ‘Love’: Evading Hate Speech Detection”. In: Proceedings of the 11th ACM Workshop on Artificial Intelligence and Security, pp. 2-12, 2018.
D. Marzoog and H. Çakir. “Deobfuscating Iraqi arabic leetspeak for hate speech detection using AraBERT and hierarchical attention network (HAN)”. Electronics (Switzerland), vol. 14, no. 21, p.4318, 2025.
N. Ejaz, F. Razi and S. Choudhury. “Towards comprehensive cyberbullying detection: A dataset incorporating aggressive texts, repetition, peerness, and intent to harm”. Computers in Human Behavior, vol. 153, p. 108123, 2024.
J. Wang, K. Fu and C. T. Lu. “SOSNet: A Graph Convolutional Network Approach to Fine-Grained Cyberbullying Detection”. In: Proceedings - 2020 IEEE International Conference on Big Data, Big Data 2020, pp. 1699-1708, 2020.
B. Ogunleye and B. Dharmaraj. “The use of a large language model for cyberbullying detection”. Analytics, vol. 2, no. 3, pp. 694- 707, 2023.
Y. Kumar, K. Huang, A. Perez, G. Yang, J. J. Li, P. Morreale, D. Kruger and R. Jiang. “Bias and cyberbullying detection and data generation using transformer artificial intelligence models and top large language models”. Electronics, vol. 13, no. 17, p. 3431, 2024.
A. G. Philipo, D. Sebastian Sarwatt, J. Ding, M. Daneshmand and H. Ning. “Assessing text classification methods for cyberbullying detection on social media platforms”. IEEE Transactions on Information Forensics and Security, vol. 20, pp. 7602-7616, 2025.
M. Abusaqer, J. Saquer and H. Shatnawi. “Efficient hate speech detection: evaluating 38 models from traditional methods to transformers”. In: ACMSE 2025 - Proceedings of the 2025 ACM Southeast Conference, Association for Computing Machinery Inc., New York, pp. 203-213, 2025.
K. Gutiérrez-Batista, J. Gómez-Sánchez and C. Fernandez-Basso. “Improving automatic cyberbullying detection in social network environments by fine-tuning a pre-trained sentence transformer language model”. Social Network Analysis and Mining, vol. 14, no. 1, p. 136, 2024.
W. Sharif, S. Abdullah, S. Iftikhar, D. Al-Madani and S. Mumtaz. “Enhancing hate speech detection in the digital age: A novel model fusion approach leveraging a comprehensive dataset”. IEEE Access, vol. 12, pp. 27225-27236, 2024.
J. Fattahi, F. Sghaier, M. Mejri, S. Bahroun, R. Ghayoula and E. Manai. “Cyberbullying detection using bag-of-words, TF-IDF, parallel CNNs and BiLSTM neural networks”. Frontiers in Artificial Intelligence and Applications, vol. 389, pp. 72-83, 2024.
E. Alikhashashneh, A. Almomani, K. M. Nahar, N. Shatnawi, A. Almomani, M. Alauthman, S. Bansal and S. H. Pan. “Unified transformer framework for automated cyberbullying detection”. International Journal of Cloud Applications and Computing, vol. 15, no. 1, pp. 1-29, 2025.
T. H. H. Aldhyani, M. H. Al-Adhaileh and S. N. Alsubari. “Cyberbullying identification system based deep learning algorithms”. Electronics, vol. 11, no. 20, p. 3273, 2022.
A. F. Alqahtani and M. Ilyas. “An ensemble-based multi-classification machine learning classifiers approach to detect multiple classes of cyberbullying”. Machine Learning and Knowledge Extraction, vol. 6, no. 1, pp. 156-170, 2024.
T. Ahmed, S. Ivan, M. Kabir, H. Mahmud and K. Hasan. “Performance analysis of transformer-based architectures and their ensembles to detect trait-based cyberbullying”. Social Network Analysis and Mining, vol. 12, no. 1, p. 99, 2022.
M. T. Hasan, M. A. E. Hossain, M. S. H. Mukta, A. Akter, M. Ahmed and S. Islam. “A review on deep-learning-based cyberbullying detection”. Future Internet, vol. 15, no. 5, p. 179, 2023.
Z. Mansur, N. Omar and S. Tiun. “Twitter hate speech detection: A systematic review of methods, taxonomy analysis, challenges, and opportunities”. IEEE Access, vol. 11, pp. 16226-16249, 2023.
G. Ramos, F. Batista, R. Ribeiro, P. Fialho, S. Moro, A. Fonseca, R. Guerra, P. Carvalho, C. Marques and Silva C. “A comprehensive review on automatic hate speech detection in the age of the transformer”. Social Network Analysis and Mining, vol. 14, no. 1, p. 204, 2024.
F. Barbieri, J. Camacho-Collados, L. Neves and L. Espinosa-Anke. “TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification”. In: Findings of the Association for Computational Linguistics Findings of ACL: EMNLP 2020. pp. 1644-1650, 2020.
T. Caselli, V. Basile, J. Mitrović and M. Granitzer. “HateBERT: Retraining BERT for Abusive Language Detection in English”. In: WOAH 2021 - 5th Workshop on Online Abuse and Harms, Proceedings of the Workshop, pp. 17-25, 2021.
B. Mathew, P. Saha, S. M. Yimam, C. Biemann, P. Goyal and A. Mukherjee. “HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection”. In: 35th AAAI Conference on Artificial Intelligence, AAAI 2021, vol. 17A, pp. 14867-14875, 2020.
J. Devlin, M. W. Chang, K. Lee and K. Toutanova. “BERT: Pre- Training of Deep Bidirectional Transformers for Language Understanding”. In: Proceedings of the 2019 Conference of the North, pp. 4171-4186, 2019.
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer and V. Stoyanov. “RoBERTa: A Robustly Optimized BERT Pretraining Approach”. 2019. Available from: https://arxiv. org/abs/1907.11692 [Last accessed on 2026 Jun 17].
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens and Z. Wojna. “Rethinking the Inception Architecture for Computer Vision”. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, vol. 2016, pp. 2818- 2826, 2016.
A. Muminovic. “Moderating harm: Benchmarking large language models for cyberbullying detection in YouTube comments”. International Journal of Computer Applications, vol. 187, no. 25, pp. 1-9, 2025.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Goran Fadhil Hassan, Omar Younis Abdulhameed

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
