Implementasi Model Hybrid IndoBERT - LinearSVC untuk Deteksi Spam Judol Obfuscated pada Komentar YouTube

Authors

  • Danang Budiman Hidayat Universitas Muria Kudus, Indonesia
  • Ahmad Abdul Chamid Universitas Muria Kudus, Indonesia
  • Ahmad Jazuli Universitas Muria Kudus., Indonesia

DOI:

https://doi.org/10.35889/progresif.v22i3.3914

Keywords:

Deteksi Spam, IndoBERT, Judi Online, Karakter N-Gram, LinearSVC

Abstract

Online gambling (judol) spam comments on Indonesian YouTube are disguised using obfuscated text techniques, substituting non-standard Unicode characters designed to bypass string-matching moderation systems. Vocabulary-based Natural Language Processing (NLP) models such as IndoBERT risk representation degradation as tokenizers map non-standard characters to unknown ([UNK]) tokens. This study implemented a hybrid model using a Dual-Track Preprocessing architecture combining IndoBERT [CLS] vectors from normalized text with TF-IDF character N-Gram LinearSVC features from raw Unicode text via sparse matrix concatenation. Experiments on 6,000 YouTube judol comments show the hybrid model achieving Accuracy=0.9978, Precision=0.9956, Recall=1.0000, F1-Score=0.9978 on the test set, identical to standalone IndoBERT and outperforming standalone LinearSVC (F1=0.9944). The hybrid model preserved IndoBERT peak performance while fully closing the 4-comment obfuscated-spam gap missed by LinearSVC, with no performance penalty from fusion.

Keywords: Character N-Gram; Gambling Spam Detection; Hybrid Model; IndoBERT; LinearSVC

Abstrak

Komentar spam judi online (judol) di YouTube Indonesia disamarkan menggunakan teknik obfuscated text, yakni substitusi karakter Unicode non-standar yang dirancang untuk menghindari sistem moderasi berbasis pencocokan string. Model Natural Language Processing (NLP) berbasis kosakata seperti IndoBERT berisiko mengalami degradasi representasi karena tokenizer memetakan karakter non-standar ke token tidak dikenal ([UNK]). Penelitian ini mengimplementasikan model hybrid dengan arsitektur Dual-Track Preprocessing yang menggabungkan vektor [CLS] IndoBERT dari teks ternormalisasi dan fitur TF-IDF karakter N-Gram LinearSVC dari teks Unicode mentah melalui konkatenasi sparse matrix. Eksperimen pada 6.000 sampel komentar judol YouTube menunjukkan model hybrid mencapai Accuracy=0,9978, Precision=0,9956, Recall=1,0000, F1-Score=0,9978 pada test set, identik dengan IndoBERT standalone dan unggul atas LinearSVC standalone (F1=0,9944). Model hybrid mempertahankan kinerja puncak IndoBERT sekaligus menutup seluruh celah 4 komentar obfuscated yang terlewat oleh LinearSVC, tanpa penalti performa dari proses fusi.

Author Biographies

Danang Budiman Hidayat, Universitas Muria Kudus

Teknik Informatika

Ahmad Abdul Chamid, Universitas Muria Kudus

Teknik Informatika

Ahmad Jazuli, Universitas Muria Kudus.

Teknik Informatika

References

Pusat Pelaporan dan Analisis Transaksi Keuangan (PPATK), "Laporan Tahunan PPATK 2023," PPATK, Jakarta, 2023. [Online]. Tersedia: https://www.ppatk.go.id

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in Proc. 2019 Conf. North American Chapter ACL: Human Language Technologies (NAACL-HLT), Minneapolis, MN, 2019, pp. 4171-4186, doi: 10.18653/v1/N19-1423.

B. Wilie et al., "IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding," in Proc. 1st Conf. Asia-Pacific Chapter ACL 8th Int. Joint Conf. NLP (AACL-IJCNLP 2020), Suzhou, China, Dec. 2020, pp. 843-857. [Online]. Available: https://aclanthology.org/2020.aacl-main.85

R. B. Perdana, Ardin, I. Budi, A. B. Santoso, A. Ramadiah, and P. K. Putra, "Detecting Online Gambling Promotions on Indonesian Twitter Using Text Mining Algorithm," Int. J. Adv. Comput. Sci. Appl. (IJACSA), vol. 15, no. 8, 2024, doi: 10.14569/IJACSA.2024.0150893.

K. Kamdan, M. P. Anugrah, M. J. Almutaali, R. Ramdani, and I. L. Kharisma, "Performance Analysis of IndoBERT for Detection of Online Gambling Promotion in YouTube Comments," in Proc. 7th Int. Global Conf. Ser. ICT Integration Technical Education & Smart Society, MDPI, Sep. 2025, p. 66, doi: 10.3390/engproc2025107066.

S. P. Santika, M. A. Saputra, A. Zahra, and D. Suhartono, "Online Gambling Promotion Detection in Indonesian YouTube Comments Using Semi-Supervised IndoBERT Classification," Procedia Comput. Sci., vol. 269, pp. 1269-1278, 2025, doi: 10.1016/j.procs.2025.09.068.

S. A. Nugraha, C. Lestari, K. B. Sanjaya, R. A. Naya, and J. Jolie, "Comparative Analysis of IndoBERT and mBERT for Online Gambling Comment Detection in Indonesian Social Media," J. Tek. Inform. (Jutif), vol. 7, no. 2, pp. 1931-1943, Apr. 2026, doi: 10.52436/1.jutif.2026.7.2.5677.

M. C. T. Manullang, A. Z. Rakhman, H. Tantriawan, and A. Setiawan, "Comparative Analysis of CNN, Transformers, and Traditional ML for Classifying Online Gambling Spam Comments in Indonesian," J. Appl. Inform. Comput. (JAIC), vol. 9, no. 3, pp. 592-602, 2025, doi: 10.30871/jaic.v9i3.9468.

F. M. Al Simabua and L. Alfat, "Optimization of IndoBERT-Lite Fine-Tuning for Spam Detection in Digital Customer Services," Sistemasi: J. Sist. Inf., vol. 15, no. 5, pp. 1886-1899, 2026, doi: 10.32520/stmsi.v15i5.6398.

M. B. M. Amin, G. Hakim, M. T. Maulana, M. F. Alwan, H. S. Anggraheni, M. J. Naufal, and N. Yudistira, "Deteksi Spam Berbahasa Indonesia berbasis Teks menggunakan Model Bert," J. Teknol. Inf. Dan Ilmu Komput. (JTIIK), vol. 11, no. 6, pp. 1291-1302, Des. 2024, doi: 10.25126/jtiik.2024118121.

A. Ghourabi and M. Alohaly, "Enhancing Spam Message Classification and Detection Using Transformer-Based Embedding and Ensemble Learning," Sensors, vol. 23, no. 8, pp. 1-17, 2023, doi: 10.3390/s23083861.

N. Arifin, U. Enri, and N. Sulistiyowati, "Penerapan Algoritma Support Vector Machine (SVM) dengan TF-IDF N-Gram untuk Text Classification," STRING (Satuan Tulisan Ris. dan Inov. Teknol.), vol. 6, no. 2, p. 129, 2021, doi: 10.30998/string.v6i2.10133.

A. Gasparetto, M. Marcuzzo, A. Zangari, and A. Albarelli, "A Survey on Text Classification Algorithms: From Text to Predictions," Information, vol. 13, no. 2, p. 83, Feb. 2022, doi: 10.3390/INFO13020083.

M. Y. Ridho and E. Yulianti, "From Text to Truth: Leveraging IndoBERT and Machine Learning Models for Hoax Detection in Indonesian News," J. Ilm. Tek. Elektro Komput. dan Inform. (JITEKI), vol. 10, no. 3, pp. 544-555, 2024, doi: 10.26555/jiteki.v10i3.29450.

H. Murfi, S. T. Gowandi, G. Ardaneswari, and S. Nurrohmah, "BERT-based Combination of Convolutional and Recurrent Neural Network for Indonesian Sentiment Analysis," Appl. Soft Comput., vol. 151, p. 111112, Jan. 2024, doi: 10.1016/j.asoc.2023.111112.

R. I. Firdaus, A. A. Chamid, and A. Jazuli, "Penerapan Algoritma Support Vector Machine dalam Analisis Sentimen terhadap Ulasan Pengguna pada Aplikasi NewSakpole," Jutisi: J. Ilm. Tek. Inform. dan Sist. Inf., vol. 15, no. 2, pp. 525-536, Apr. 2026, doi: 10.35889/jutisi.v15i2.3478.

F. Alvin and N. A. S. Winarsih, "Perbandingan Kinerja Model IndoBERT, IndoBERTweet, dan Algoritma Klasik pada Analisis Sentimen Isu Indonesia Gelap," Build. Informatics Technol. Sci. (BITS), vol. 7, no. 3, pp. 1601-1613, Dec. 2025, doi: 10.47065/bits.v7i3.8636.

P. M. S. Ardinata, A. A. J. Permana, and I. N. S. W. Wijaya, "Identifikasi dan Normalisasi Teks Slang dengan FastText pada Twitter dalam Bahasa Indonesia," J. Pendidik. Teknol. dan Kejuru., vol. 21, no. 1, pp. 34-44, 2024, doi: 10.23887/jptkundiksha.v21i1.66381.

R. M. Yazid, F. R. Umbara, and P. N. Sabrina, "Deteksi Ujaran Kebencian dengan Metode Klasifikasi Naïve Bayes dan Metode N-Gram pada Dataset Multi-Label Twitter Berbahasa Indonesia," Informatics and Digital Expert (INDEX), vol. 4, no. 2, pp. 46-52, 2023, doi: 10.36423/index.v4i2.894.

M. V. S. Handayani and Muljono, "Arsitektur Hibrida IndoBERTweet-CNN untuk Klasifikasi Ujaran Kebencian Berbahasa Gaul di Media Sosial," Infotekmesin, vol. 17, no. 1, pp. 39-47, Jan. 2026, doi: 10.35970/infotekmesin.v17i1.3026.

A. N. Khoruzhaya, D. V. Kozlov, K. M. Arzamasov, and E. I. Kremneva, "Comparison of an Ensemble of Machine Learning Models and the BERT Language Model for Analysis of Text Descriptions of Brain CT Reports to Determine the Presence of Intracranial Hemorrhage," Sovremennye Tehnologii v Medicine, vol. 16, no. 1, p. 27, 2024, doi: 10.17691/stm2024.16.1.03.

M. Putri, M. Afdal, R. Novita, and M. Mustakim, "Perbandingan Evaluasi Kernel Support Vector Machine dalam Analisis Sentimen Chatbot AI pada Ulasan Google Play Store," J. Teknol. Sist. Inf. dan Apl. (JTSI), vol. 7, no. 3, pp. 1236-1245, Jul. 2024, doi: 10.32493/jtsi.v7i3.41354.

N. A. Nevrada and M. A. Syaputra, "Sentiment Analysis of Telegram App Reviews on Google Play Store Using the Support Vector Machine (SVM) Algorithm," J. Appl. Inform. Comput. (JAIC), vol. 9, no. 1, pp. 96-105, Jan. 2025, doi: 10.30871/jaic.v9i1.8851.

S. A. S. Mola, D. L. B. Baun, I. O. Nunes, and M. M. A. R. Sani, "Analisis Sentimen Aplikasi Halo BCA di Google Play Store Menggunakan Metode Naive Bayes, Support Vector Machine, dan Random Forest," HOAQ: High Educ. Organ. Arch. Qual. – J. Teknol. Inf., vol. 15, no. 2, pp. 69-79, Dec. 2024, doi: 10.52972/hoaq.vol15no2.p69-79.

P. Arsi and R. Waluyo, "Analisis Sentimen Wacana Pemindahan Ibu Kota Indonesia Menggunakan Algoritma Support Vector Machine (SVM)," J. Teknol. Inf. dan Ilmu Komput. (JTIIK), vol. 8, no. 1, p. 147, 2021, doi: 10.25126/jtiik.0813944.

A. A. Chamid, R. Nindyasari, N. Azizah, and A. Hariyadi, "Analysis of Public Opinion on the Governor Candidate Debate Using LDA and IndoBERT," Kinetik: Game Technol. Inf. Syst. Comput. Netw. Comput. Electron. Control, vol. 10, no. 3, Aug. 2025, doi: 10.22219/kinetik.v10i3.2221.

A. A. Chamid, Widowati, and R. Kusumaningrum, "Graph-Based Semi-Supervised Deep Learning for Indonesian Aspect-Based Sentiment Analysis," Big Data Cogn. Comput., vol. 7, no. 1, p. 5, Mar. 2023, doi: 10.3390/bdcc7010005.

A. Jazuli, Widowati, A. A. Chamid, and R. Kusumaningrum, "Transformer-Based Semantic Indexing for Aspect-Based Sentiment Analysis Using an Enhanced Index Generation Algorithm with BERT," Int. J. Adv. Technol. Eng. Explor. (IJATEE), vol. 12, no. 127, pp. 907-926, 2025, doi: 10.19101/IJATEE.2024.111102114.

A. Jazuli, Widowati, and R. Kusumaningrum, "Optimizing Aspect-Based Sentiment Analysis Using BERT for Comprehensive Analysis of Indonesian Student Feedback," Appl. Sci., vol. 15, no. 1, pp. 1-28, 2025, doi: 10.3390/app15010172.

A. A. Chamid, Widowati, and R. Kusumaningrum, "Labeling Consistency Test of Multi-Label Data for Aspect and Sentiment Classification Using the Cohen Kappa Method," Ingénierie des Systèmes d’Information, vol. 29, no. 1, pp. 161-167, 2024, doi: 10.18280/isi.290118.

A. A. Chamid, W. Widowati, and R. Kusumaningrum, "Text Data Labeling Process for Semi-Supervised Learning Modeling," in 12th Int. Semin. New Paradigm Innov. Nat. Sci. Its Appl. (12th ISNPINSA), AIP Conf. Proc., vol. 3165, p. 030011, 2024, doi: 10.1063/5.0216320.

A. A. Chamid, Widowati, and R. Kusumaningrum, "Multi-Label Text Classification on Indonesian User Reviews Using Semi-Supervised Graph Neural Networks," ICIC Express Lett., vol. 17, no. 10, pp. 1075-1084, 2023, doi: 10.24507/

Downloads

Published

2026-07-15

How to Cite

Budiman Hidayat, D., Abdul Chamid, A., & Jazuli, A. (2026). Implementasi Model Hybrid IndoBERT - LinearSVC untuk Deteksi Spam Judol Obfuscated pada Komentar YouTube. Progresif: Jurnal Ilmiah Komputer, 22(3), 820–833. https://doi.org/10.35889/progresif.v22i3.3914

Issue

Section

Articles

Citation Check

Similar Articles

<< < 1 2 3 4 5 6 7 8 9 > >> 

You may also start an advanced similarity search for this article.