Klasifikasi Named Entity Recognition Pada Cerita Pendek Bahasa Indonesia Menggunakan Model Indobert

Authors

  • Nazifa Samsurizal Universitas Islam Negeri Ar-Raniry, Indonesia
  • Hendri Ahmadian UIN Ar-Raniry Banda Aceh, Indonesia
  • Nurrizqa Nurrizqa UIN Ar-Raniry Banda Aceh, Indonesia

DOI:

https://doi.org/10.35889/progresif.v22i3.3731

Keywords:

Named Entity Recognition, IndoBERT, cerita pendek bahasa Indonesia, Natural Language Processing, Transformer

Abstract

Entity recognition in Indonesian short stories requires a model capable of understanding narrative language context effectively. This study aimed to implement the IndoBERT model for the Named Entity Recognition (NER) task to identify Person, Location, and Organization entities. The dataset was obtained from the Majalah Bobo e-book and organized using the BIO (Begin, Inside, Outside) format. The research included pre-processing, tokenization, label encoding, token alignment, and fine-tuning using the AutoModelForTokenClassification model. Evaluation was conducted using precision, recall, and F1-score metrics. The results showed that IndoBERT achieved F1-scores of 0.95 for Person, 0.81 for Organization, and 0.75 for Location, with a weighted average F1-score of 0.90. These results indicated that IndoBERT performed entity recognition effectively on Indonesian short stories.

Keywords: Named Entity Recognition; IndoBERT; Indonesian short stories; Natural Language Processing; Transformer. 

 

Abstrak

Pengenalan entitas pada cerita pendek bahasa Indonesia memerlukan model yang mampu memahami konteks bahasa naratif secara efektif. Penelitian ini bertujuan menerapkan model IndoBERT pada tugas Named Entity Recognition (NER) untuk mengenali entitas Person, Location, dan Organization. Dataset diperoleh dari e-book Majalah Bobo dan disusun menggunakan format BIO (Begin, Inside, Outside). Penelitian dilakukan melalui tahapan pre-processing, tokenisasi, label encoding, token alignment, dan fine-tuning menggunakan model AutoModelForTokenClassification. Evaluasi dilakukan menggunakan metrik precision, recall, dan F1-score. Hasil pengujian menunjukkan bahwa model IndoBERT memperoleh nilai F1-score sebesar 0.95 pada label Person, 0.81 pada Organization, dan 0.75 pada Location, dengan weighted average F1-score sebesar 0.90. Hasil tersebut menunjukkan bahwa IndoBERT mampu melakukan pengenalan entitas dengan baik pada cerita pendek bahasa Indonesia.

Kata kunci: Named Entity Recognition; IndoBERT; cerita pendek bahasa Indonesia; Natural Language Processing; Transformer.

Author Biographies

Nazifa Samsurizal, Universitas Islam Negeri Ar-Raniry

Teknologi Informasi

Hendri Ahmadian, UIN Ar-Raniry Banda Aceh

Teknologi Informasi

Nurrizqa Nurrizqa, UIN Ar-Raniry Banda Aceh

Teknologi Informasi

References

[1] M. Amien and G. F. Gunawan, “BERT dan Bahasa Indonesia: Studi tentang Efektivitas Model NLP Berbasis Transformer,” ELANG: Journal of Interdisciplinary Research, vol. 1, no. 02, pp. 132–140, 2024.

[2] A. F. Muhammad and M. S. Hasibuan, “View of Peningkatan Akurasi Named Entity Recognition (NER) Dengan Fine-Tuning BERT Pada Dataset Bahasa Indonesia,” CESS (Journal of Computer Engineering, System and Science), vol. 10, no. 02, pp. 702–713, 2025.

[3] B. Jehangir, S. Radhakrishnan, and R. Agarwal, “A survey on Named Entity Recognition—datasets, tools, and methodologies. Natural Language Processing Journal, 3, 100017, pp. 1–12,” 2023.

[4] M. T. A. Anwar, S. H. Wijoyo, and W. H. N. Putra, “Implementasi Metode TextRank dan Named Entity Recognition Untuk Ekstraksi Kata Kunci Pada Media Online Berita,” Jurnal Sistem Informasi, Teknologi Informasi, dan Edukasi Sistem Informasi, vol. 5, no. 1, pp. 34–41, 2024.

[5] D. Agustini, M. I. Firdaus, M. Farida, M. E. Rosadi, H. Noor, and R. Muttaqin, “Peningkatan Kinerja Named Entity Recognition Bahasa Indonesia Melalui Augmentasi Data Berbasis Large Language Models,” Jurnal Informatika Teknologi dan Sains (Jinteks), vol. 7, no. 3, pp. 1361–1369, 2025.

[6] N. U. Muchtar, D. Arifianto, and R. Umilasari, “Analisis Kinerja Transformer Untuk Named Entity Recognition (NER) Menggunakan IndoBERT Pada Transkrip Video Politik Berbahasa Indonesia,” Discovery: Jurnal Ilmu Pengetahuan, vol. 10, no. 2, pp. 180–199, 2025.

[7] S. Dharmawan, V. C. Mawardi, and N. J. Perdana, “Klasifikasi Ujaran Kebencian Menggunakan Metode FeedForward Neural Network (IndoBERT),” Jurnal Ilmu Komputer dan Sistem Informasi, vol. 11, no. 1, pp. 1–6, 2023.

[8] M. Amien, “Sejarah dan perkembangan teknik Natural Language Processing (NLP) bahasa Indonesia: Tinjauan tentang sejarah, perkembangan teknologi, dan aplikasi NLP dalam bahasa Indonesia,” arXiv preprint arXiv:2304.02746, pp. 1–7, 2023.

[9] Y. O. Sihombing and N. V. Situmorang, “Prediksi Sentimen Pada Teks Media Sosial Corporate University Menggunakan RoBERTa,” Prosiding PITNAS Widyaiswara, vol. 1, pp. 302–316, 2024.

[10] E. Yulianti, N. Bhary, J. Abdurrohman, F. W. Dwitilas, E. Q. Nuranti, and H. S. Husin, “Named entity recognition on Indonesian legal documents: a dataset and study using transformer-based models.,” International Journal of Electrical & Computer Engineering (2088-8708), vol. 14, no. 5, pp. 5489–5501, 2024.

[11] N. R. Jonathan, Syafrijon, D. Novaliendry, and K. Budayawan, “View of Ekstraksi Entitas Keterampilan pada Teks Lowongan Pekerjaan Berbahasa Indonesia Menggunakan Model IndoBERT,” Jurnal Pendidikan Tambusai, vol. 9, no. 3, pp. 40145–40171, 2025.

[12] W. L. Seow, I. Chaturvedi, A. Hogarth, R. Mao, and E. Cambria, “A review of named entity recognition: from learning methods to modelling paradigms and tasks,” Artif. Intell. Rev., vol. 58, no. 10, p. 315, 2025.

[13] S. Khairina, N. Saffa, D. C. U. Lieharyani, and J. Hutahaean, “Contextualized Word Embedding Untuk Ekstraksi Kutipan Berita Indonesia,” Jurnal Ilmiah Matrik, vol. 27, no. 2, pp. 142–153, 2025.

[14] I. I. Sholikhah, A. T. J. Harjanta, and K. Latifah, “Machine Learning Untuk Deteksi Berita Hoax Menggunakan BERT,” in Prosiding Seminar Nasional Informatika, 2023, pp. 524–531.

[15] S. C. Kusuma, T. N. Fatyanosa, and M. Data, “Fine-Tuning Bert Untuk Named Entity Recognition Pada Dokumen Klinis Gizi,” Jurnal Pengembangan Teknologi Informasi dan Ilmu Komputer, vol. 10, no. 2, pp. 1–10, 2026.

[16] J. Opitz, “A closer look at classification evaluation metrics and a critical reflection of common evaluation practice,” Trans. Assoc. Comput. Linguist., vol. 12, pp. 820–836, 2024.

[17] P. Srinivasan and R. Venkatakrishnan, “Transformer-based models for named entity recognition: A comparative study,” in 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT), IEEE, 2023, pp. 1–5.

[18] N. Patwardhan, S. Marrone, and C. Sansone, “Transformers in the real world: A survey on nlp applications,” Information, vol. 14, no. 4, p. 242, 2023.

[19] I. Keraghel, S. Morbieu, and M. Nadif, “Recent advances in named entity recognition: A comprehensive survey and comparative study,” arXiv preprint arXiv:2401.10825, pp. 1–42, 2024.

Downloads

Published

2026-07-14

How to Cite

Samsurizal, N., Ahmadian, H., & Nurrizqa, N. (2026). Klasifikasi Named Entity Recognition Pada Cerita Pendek Bahasa Indonesia Menggunakan Model Indobert. Progresif: Jurnal Ilmiah Komputer, 22(3), 707–716. https://doi.org/10.35889/progresif.v22i3.3731

Issue

Section

Articles

Citation Check

Similar Articles

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 > >> 

You may also start an advanced similarity search for this article.