Perancangan Failure-Aware Self-Healing Workflow pada Sistem Event-Driven untuk Pemulihan Proses Layanan Digital Secara Otomatis

Authors

  • Dimaz Ardawan Universitas Negeri Padang, Indonesia
  • Denny Kurniadi Universitas Negeri Padang, Indonesia
  • Khairi Budayawan Universitas Negeri Padang, Indonesia
  • Randi Proska Sandra Universitas Negeri Padang, Indonesia

DOI:

https://doi.org/10.35889/progresif.v22i3.3977

Keywords:

Failure-Aware, Self-Healing, Arsitektur Event-Driven, Orkestrasi Workflow, Mean Time to Recovery

Abstract

Digital service systems built on event-driven architecture relied on infrastructure-level self-healing, leaving workflow-layer failures handled reactively and processes abandoned mid-execution even when infrastructure remained available. This study designed a failure-aware self-healing workflow that detected, classified, and recovered digital service failures automatically based on event context, treating classification as mandatory before recovery. A prototype was built following the Prototyping Model, integrating n8n as workflow orchestrator and Apache Kafka as event broker inside Docker, with recovery governed by a JavaScript finite state machine, tested on five failure scenarios and two baseline conditions, each run for 30 iterations. The system classified all failures correctly across 150 iterations, reaching 100% recovery accuracy and a 100% success rate where fallback was available. Self-healing lowered Recovery Time by 29.90% against a no-recovery baseline, though a static-recovery baseline reached a lower time with 0% success, showing recovery speed and correctness can move in opposite directions.

Keywords: Failure-Aware; Self-Healing; Event-Driven Architecture; Workflow Orchestration; Mean Time to Recovery (MTTR)

Abstrak

Sistem layanan digital berbasis arsitektur event-driven umumnya mengandalkan self-healing pada tingkat infrastruktur, sementara kegagalan pada lapisan orkestrasi workflow ditangani secara reaktif melalui pemantauan manual sehingga proses bisnis dapat terhenti di tengah eksekusi meskipun infrastrukturnya masih berjalan normal. Penelitian ini merancang failure-aware self-healing workflow yang mendeteksi, mengklasifikasikan, dan memulihkan kegagalan layanan digital secara otomatis berdasarkan konteks event, dengan klasifikasi sebagai tahap wajib sebelum pemulihan dijalankan. Prototipe dibangun mengikuti Prototyping Model, mengintegrasikan n8n sebagai orkestrator workflow dan Apache Kafka sebagai event broker di dalam Docker, dengan keputusan pemulihan dikendalikan finite state machine berbasis JavaScript, diuji pada lima skenario kegagalan dan dua baseline, masing-masing 30 iterasi. Sistem mengklasifikasikan seluruh kegagalan secara tepat pada 150 iterasi, mencapai recovery accuracy 100% dan success rate 100% pada skenario dengan jalur fallback. Self-healing menurunkan MTTR sebesar 29,90% dibandingkan baseline tanpa pemulihan, meskipun baseline pemulihan statis mencatat waktu lebih rendah namun dengan success rate 0%, menunjukkan kecepatan dan ketepatan pemulihan dapat bergerak berlawanan arah.

Author Biographies

Dimaz Ardawan, Universitas Negeri Padang

Informatika

Denny Kurniadi, Universitas Negeri Padang

Informatika

Khairi Budayawan, Universitas Negeri Padang

Informatika

Randi Proska Sandra, Universitas Negeri Padang

Informatika

References

[1] S. Kul, I. Tashiev, A. Şentaş, and A. Sayar, “Event-Based Microservices With Apache Kafka Streams: A Real-time Vehicle Detection System Based on Type, Color, and Speed Attributes,” IEEE Access, vol. 9, pp. 83137–83148, 2021, doi: 10.1109/ACCESS.2021.3085736.

[2] X. Zhou et al., “Fault Analysis and Debugging of Microservice Systems: Industrial Survey, Benchmark System, and Empirical Study,” IEEE Transactions on Software Engineering, vol. 47, no. 2, pp. 243–260, 2021, doi: 10.1109/TSE.2018.2887384.

[3] A. Alnafessah, A. U. Gias, R. Wang, L. Zhu, G. Casale, and A. Filieri, “Quality-Aware DevOps Research: Where Do We Stand?,” IEEE Access, vol. 9, pp. 44476–44489, 2021, doi: 10.1109/ACCESS.2021.3064867.

[4] P. Desaraju, “Self-healing Software Systems: AI-Driven Fault Prediction and Recovery,” World Journal of Advanced Engineering Technology and Sciences, vol. 18, no. 3, pp. 064–071, Mar. 2026, doi: 10.30574/wjaets.2026.18.3.0120.

[5] B. Magableh and M. Almiani, “A Self Healing Microservices Architecture: A Case Study in Docker Swarm Cluster,” in Advances in Intelligent Systems and Computing, Springer Verlag, 2020, pp. 846–858. doi: 10.1007/978-3-030-15032-7_71.

[6] J. A. Rasheedh, G. N. Ahmed, S. H. Abdul Cader, and K. Nirmala, “Fault tolerance Enhancement in Microservices using Orchestration Process with Agile,” in 2024 8th International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC), 2024, pp. 406–413. doi: 10.1109/I-SMAC61858.2024.10714589.

[7] J. Nikolic, N. Jubatyrov, and E. Pournaras, “Self-healing Dilemmas in Distributed Systems: Fault correction vs. Fault tolerance,” IEEE Transactions on Network and Service Management, vol. 18, no. 3, pp. 2728–2741, Sep. 2021, doi: 10.1109/TNSM.2021.3092939.

[8] A. Bucchiarone, C. Guidi, I. Lanese, N. Bencomo, and J. Spillner, “A MAPE-K Approach to Autonomic Microservices,” in 2022 IEEE 19th International Conference on Software Architecture Companion (ICSA-C), 2022, pp. 100–103. doi: 10.1109/ICSA-C54293.2022.00025.

[9] T. Wang, W. Zhang, J. Xu, and Z. Gu, “Workflow-Aware Automatic Fault Diagnosis for Microservice-Based Applications With Statistics,” IEEE Transactions on Network and Service Management, vol. 17, no. 4, pp. 2350–2363, 2020, doi: 10.1109/TNSM.2020.3022028.

[10] Y. Pan, M. Ma, X. Jiang, and P. Wang, “DyCause: Crowdsourcing to Diagnose Microservice Kernel Failure,” IEEE Trans. Dependable Secure Comput., vol. 20, no. 6, pp. 4763–4777, 2023, doi: 10.1109/TDSC.2022.3233915.

[11] R. S. Pressman, Software engineering: A practitioner’s approach, 7th ed. McGraw-Hill, 2010.

[12] A. Al-Said Ahmad, L. F. Al-Qora’n, and A. Zayed, “Exploring the impact of chaos engineering with various user loads on cloud native applications: an exploratory empirical study,” Computing, vol. 106, no. 7, pp. 2389–2425, Jul. 2024, doi: 10.1007/s00607-024-01292-z.

[13] n8n, “n8n Documentation,” docs.n8n.io. Accessed: May 06, 2026. [Online]. Available: [https://docs.n8n.io/]

[14] Apache Software Foundation, “Apache Kafka Documentation,” kafka.apache.org. Accessed: May 06, 2026. [Online]. Available: [https://kafka.apache.org/documentation/]

[15] S. Aydin and C. B. Çebi, “Comparison of Choreography vs Orchestration Based Saga Patterns in Microservices,” in 2022 International Conference on Electrical, Computer and Energy Technologies (ICECET), 2022, pp. 1–6. doi: 10.1109/ICECET55527.2022.9872665.

[16] H. Zhang, A. Cardoza, P. B. Chen, S. Angel, and V. Liu, “Fault-tolerant and transactional stateful serverless workflows,” in 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20), USENIX Association, Nov. 2020, pp. 1187–1204. [Online]. Available: https://www.usenix.org/conference/osdi20/presentation/zhang-haoran

[17] T. P. Raptis, C. Cicconetti, and A. Passarella, “Efficient topic partitioning of Apache Kafka for high-Reliability real-time data streaming applications,” Future Generation Computer Systems, vol. 154, pp. 173–188, May 2024, doi: 10.1016/j.future.2023.12.028.

Downloads

Published

2026-07-15

How to Cite

Ardawan, D., Kurniadi, D., Budayawan, K., & Proska Sandra, R. (2026). Perancangan Failure-Aware Self-Healing Workflow pada Sistem Event-Driven untuk Pemulihan Proses Layanan Digital Secara Otomatis. Progresif: Jurnal Ilmiah Komputer, 22(3), 779–793. https://doi.org/10.35889/progresif.v22i3.3977

Issue

Section

Articles

Citation Check

Similar Articles

1 2 3 4 5 6 7 8 9 10 11 12 13 > >> 

You may also start an advanced similarity search for this article.