Perbandingan Prediksi Loan Default Menggunakan Pemodelan Algoritma Random Forest dan XGBoost untuk Mitigasi Risiko Kredit Bank
(Studi Kasus: BTN Syariah)
DOI:
https://doi.org/10.55826/jtmit.v5i3.2101Keywords:
gagal bayar, non performing financing, early warning system, risiko kredit, random forest, xgboostAbstract
Peningkatan penyaluran pembiayaan syariah pascapandemi meningkatkan risiko gagal bayar yang mengancam stabilitas bank. Bank BTN Syariah Cabang X mencatat Non Performing Financing (NPF) sebesar 15,27% pada 2019, turun ke 1,47% pada 2022, namun meningkat kembali pada 2023-2024. Penelitian ini membangun dan membandingkan model prediksi gagal bayar pembiayaan konsumer menggunakan algoritma Random Forest dengan Synthetic Minority Over-sampling Technique (SMOTE) dan Extreme Gradient Boosting (XGBoost) dengan Weighted Class sebagai fondasi Early Warning System (EWS) untuk monitoring nasabah eksisting secara periodik. Data yang digunakan adalah 14.374 nasabah pembiayaan konsumer aktif per 31 Desember 2024 dengan rasio ketidakseimbangan kelas sebesar 77,5:1. Model dievaluasi menggunakan recall, precision, F1-Score, ROC-AUC, dan PR-AUC dengan penekanan pada recall sebagai metrik prioritas. Hasil menunjukkan bahwa Random Forest dengan SMOTE merupakan model terbaik dengan recall kelas default = 0,6757 dan ROC-AUC = 0,9368, mendeteksi 25 dari 37 nasabah gagal bayar pada data testing. XGBoost menghasilkan F1-Score = 0,4878 dan PR-AUC = 0,4849 lebih tinggi namun recall lebih rendah (0,5405). Analisis feature importance menggunakan MDI dan SHAP mengungkapkan bahwa Saldo Tersedia dan Rasio Saldo/Angsuran merupakan prediktor dominan, membuktikan likuiditas real time sebagai leading indicator yang lebih kuat dari riwayat tunggakan.
References
[1] A. Yuwannita, R. Mulyany, and H. Fahlevi, “What Causes Non-Performance Financing? Insights From Islamic Commercial Banks in Indonesia and Malaysia,” Muqtasid J. Ekon. dan Perbank. Syariah, vol. 13, no. 1, pp. 77–94, Oct. 2022, doi: 10.18326/muqtasid.v13i1.77-94.
[2] S. Anwar and A. . H. Ali, “ANNs-Based Early Warning System for Indonesian Islamic Banks,” Bul. Ekon. Monet. dan Perbank., vol. 20, no. 3, pp. 325–342, Jan. 2018, doi: 10.21098/bemp.v20i3.856.
[3] G. Chironna and G. Orlando, “Predicting bank defaults in Italy: A comparative analysis of conventional and machine learning approaches,” Econ. Anal. Policy, vol. 89, pp. 788–833, Jan. 2026, doi: 10.1016/j.eap.2025.12.002.
[4] L. Zhu, D. Qiu, D. Ergu, C. Ying, and K. Liu, “A study on predicting loan default based on the random forest algorithm,” Procedia Comput. Sci., vol. 162, pp. 503–513, 2019, doi: 10.1016/j.procs.2019.12.017.
[5] D. Lin, Y. Li, and R. Zhang, “A SHAP-Based Interpretability Analysis of XGBoost for Loan Risk Prediction,” in Proceedings of the 2025 International Conference on Information Economy, Data Modeling and Cloud Computing, New York, NY, USA: ACM, Aug. 2025, pp. 219–228. doi: 10.1145/3772900.3772936.
[6] I. Brown and C. Mues, “An experimental comparison of classification algorithms for imbalanced credit scoring data sets,” Expert Syst. Appl., vol. 39, no. 3, pp. 3446–3453, Feb. 2012, doi: 10.1016/j.eswa.2011.09.033.
[7] F. Weng, M. Zhu, M. Buckle, P. Hajek, and M. Z. Abedin, “Class imbalance Bayesian model averaging for consumer loan default prediction: The role of soft credit information,” Res. Int. Bus. Financ., vol. 74, p. 102722, Feb. 2025, doi: 10.1016/j.ribaf.2024.102722.
[8] S. Shi, R. Tse, W. Luo, S. D’Addona, and G. Pau, “Machine learning-driven credit risk: a systemic review,” Neural Comput. Appl., vol. 34, no. 17, pp. 14327–14339, Sep. 2022, doi: 10.1007/s00521-022-07472-2.
[9] F. O. Aghware et al., “Enhancing the Random Forest Model via Synthetic Minority Oversampling Technique for Credit-Card Fraud Detection,” J. Comput. Theor. Appl., vol. 1, no. 4, pp. 407–420, Mar. 2024, doi: 10.62411/jcta.10323.
[10] K. Wang, J. Wan, G. Li, and H. Sun, “A Hybrid Algorithm-Level Ensemble Model for Imbalanced Credit Default Prediction in the Energy Industry,” Energies, vol. 15, no. 14, p. 5206, Jul. 2022, doi: 10.3390/en15145206.
[11] M. Asutay and J. Othman, “Alternative measures for predicting financial distress in the case of Malaysian Islamic banks: assessing the impact of global financial crisis,” J. Islam. Account. Bus. Res., vol. 11, no. 10, pp. 1827–1845, Dec. 2020, doi: 10.1108/JIABR-12-2019-0223.
[12] Y. Chen, R. Calabrese, and B. Martin-Barragan, “Interpretable machine learning for imbalanced credit scoring datasets,” Eur. J. Oper. Res., vol. 312, no. 1, pp. 357–372, Jan. 2024, doi: 10.1016/j.ejor.2023.06.036.
[13] L. Yu, R. Zhou, R. Chen, and K. K. Lai, “Missing Data Preprocessing in Credit Classification: One-Hot Encoding or Imputation?,” Emerg. Mark. Financ. Trade, vol. 58, no. 2, pp. 472–482, Jan. 2022, doi: 10.1080/1540496X.2020.1825935.
[14] N. Alamsyah, B. Budiman, T. Parama Yoga, and R. Y. Rakhman Alamsyah, “A stacking ensemble model with SMOTE for improved imbalanced classification on credit data,” TELKOMNIKA (Telecommunication Comput. Electron. Control., vol. 22, no. 3, p. 657, Feb. 2024, doi: 10.12928/telkomnika.v22i3.25921.
[15] B. Bischl et al., “Hyperparameter optimization: Foundations, algorithms, best practices, and open challenges,” WIREs Data Min. Knowl. Discov., vol. 13, no. 2, Mar. 2023, doi: 10.1002/widm.1484.
[16] Y. Song and Y. Peng, “A MCDM-Based Evaluation Approach for Imbalanced Classification Methods in Financial Risk Prediction,” IEEE Access, vol. 7, pp. 84897–84906, 2019, doi: 10.1109/ACCESS.2019.2924923.
[17] T. Chen and C. Guestrin, “XGBoost,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, New York, NY, USA: ACM, Aug. 2016, pp. 785–794. doi: 10.1145/2939672.2939785.
[18] M. Moscatelli, F. Parlapiano, S. Narizzano, and G. Viggiano, “Corporate default forecasting with machine learning,” Expert Syst. Appl., vol. 161, p. 113567, Dec. 2020, doi: 10.1016/j.eswa.2020.113567.
[19] T. Shoko, T. Verster, and L. Dube, “Comparative analysis of classical and Bayesian optimisation techniques: Impact on model performance and interpretability in credit risk modelling using SHAP and PDPs,” Data Sci. Financ. Econ., vol. 5, no. 3, pp. 320–354, 2025, doi: 10.3934/DSFE.2025014.
[20] M. Drehmann and M. Juselius, “Evaluating early warning indicators of banking crises: Satisfying policy requirements,” Int. J. Forecast., vol. 30, no. 3, pp. 759–780, Jul. 2014, doi: 10.1016/j.ijforecast.2013.10.002.
[21] T. Saito and M. Rehmsmeier, “The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets,” PLoS One, vol. 10, no. 3, p. e0118432, Mar. 2015, doi: 10.1371/journal.pone.0118432.
[22] X. Zhu, Q. Chu, X. Song, P. Hu, and L. Peng, “Explainable prediction of loan default based on machine learning models,” Data Sci. Manag., vol. 6, no. 3, pp. 123–133, Sep. 2023, doi: 10.1016/j.dsm.2023.04.003.
[23] H. Li and W. Wu, “Loan default predictability with explainable machine learning,” Financ. Res. Lett., vol. 60, p. 104867, Feb. 2024, doi: 10.1016/j.frl.2023.104867.
[24] J. T. Hancock and T. M. Khoshgoftaar, “CatBoost for big data: an interdisciplinary review,” J. Big Data, vol. 7, no. 1, p. 94, Dec. 2020, doi: 10.1186/s40537-020-00369-8.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Gigih Prakoso Wigantiyoko, R. Mohamad Atok

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.













