Machine Learning-Based Predictive Fault Detection and Resilience Optimization in Large-Scale Distributed Systems
Keywords:
Machine Learning, Predictive Fault Detection, Distributed Systems, Fault Tolerance, Resilience Optimization.Abstract
The rapid expansion of cloud computing, high-performance computing, large-scale data centers, distributed applications, and service-oriented computing infrastructures has significantly increased the complexity of modern distributed systems. Such environments consist of large numbers of interconnected computing nodes, virtual machines, storage devices, network resources, and software services that operate collaboratively to deliver continuous and scalable computing services. However, the increasing scale, heterogeneity, workload variability, and dynamic behavior of distributed infrastructures also increase their exposure to hardware faults, software failures, node crashes, resource exhaustion, communication disruptions, and performance degradation. Traditional fault-tolerance mechanisms generally react to failures after they occur and therefore may result in service interruption, resource wastage, recovery overhead, and violations of service-level requirements.
Downloads
References
Salami, H., Saadatfar, H., Rahmani Fard, F., Shekofteh, S. K., & Deldari, H. (2010). Improving cluster computing performance based on job futurity prediction. 2010 3rd International Conference on Advanced Computer Theory and Engineering (ICACTE), Vol. 6, V6-303–V6-307. https://doi.org/10.1109/ICACTE.2010.5579820
Saadatfar, H., Fadishei, H., & Deldari, H. (2012). Predicting job failures in AuverGrid based on workload log analysis. New Generation Computing, 30(1), 73–94. https://doi.org/10.1007/s00354-012-0105-z
Chen, X., Lu, C.-D., & Pattabiraman, K. (2014). Failure analysis of jobs in compute clouds: A Google cluster case study. IEEE 25th International Symposium on Software Reliability Engineering (ISSRE), 167–177. https://doi.org/10.1109/ISSRE.2014.34
Chen, X., Lu, C.-D., & Pattabiraman, K. (2014). Failure prediction of jobs in compute clouds: A Google cluster case study. IEEE 25th International Symposium on Software Reliability Engineering Workshops (ISSREW), 341–346. https://doi.org/10.1109/ISSREW.2014.105
Soualhia, M., Khomh, F., & Tahar, S. (2015). Predicting scheduling failures in the cloud: A case study with Google clusters and Hadoop on Amazon EMR. 2015 IEEE 17th International Conference on High Performance Computing and Communications, 58–65. https://doi.org/10.1109/HPCC-CSS-ICESS.2015.170
Jaiswal, I. A. (2020). Machine learning-based predictive fault detection and resilience optimization in large-scale distributed systems. International Journal of Intelligent Systems and Applications in Engineering.
Zheng, W., Wang, Z., Huang, H., Meng, L., & Qiu, X. (2016). EHMM-CT: An online method for failure prediction in cloud computing systems. KSII Transactions on Internet and Information Systems, 10(9), 4087–4107. https://doi.org/10.3837/tiis.2016.09.004
Islam, T., & Manivannan, D. (2017). Predicting application failure in cloud: A machine learning approach. 2017 IEEE 1st International Conference on Cognitive Computing (ICCC), 24–31. https://doi.org/10.1109/IEEE.ICCC.2017.11
Pitakrat, T., Okanović, D., van Hoorn, A., & Grunske, L. (2018). Hora: Architecture-aware online failure prediction. Journal of Systems and Software, 137, 669–685. https://doi.org/10.1016/j.jss.2017.02.041
Lin, Q., Hsieh, K., Dang, Y., Zhang, H., Sui, K., Xu, Y., Lou, J.-G., Li, C., Wu, Y., Yao, R., Chintalapati, M., & Zhang, D. (2018). Predicting node failure in cloud service systems. Proceedings of the 26th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 480–490. https://doi.org/10.1145/3236024.3236060
Mariani, L., Pezzè, M., Riganelli, O., & Xin, R. (2020). Predicting failures in multi-tier distributed systems. Journal of Systems and Software, 161, 110464. https://doi.org/10.1016/j.jss.2019.110464
Gummadi, V. P. K. (2020). API design and implementation: RAML and OpenAPI specification. Journal of Electrical Systems, 16(4). https://doi.org/10.52783/jes.9329
Lu, S., Luo, B., Patel, T., Yao, Y., Tiwari, D., & Shi, W. (2020). Making disk failure predictions SMARTer! 18th USENIX Conference on File and Storage Technologies (FAST '20), 151–167.
Ahmad, W., Khan, S. A., Kim, C. H., & Kim, J.-M. (2020). Feature selection for improving failure detection in hard disk drives using a genetic algorithm and significance scores. Applied Sciences, 10(9), 3200. https://doi.org/10.3390/app10093200
Li, Z., Cheng, Q., Hsieh, K., Dang, Y., Huang, P., Singh, P., Yang, X., Lin, Q., Wu, Y., Levy, S., & Chintalapati, M. (2020). Gandalf: An intelligent, end-to-end analytics service for safe deployment in large-scale cloud infrastructure. 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI '20), 389–402.
Levy, S., Yao, R., Wu, Y., Dang, Y., Huang, P., Mu, Z., Zhao, P., Ramani, T., Govindaraju, N., Li, X., Lin, Q., Shafriri, G. L., & Chintalapati, M. (2020). Predictive and adaptive failure mitigation to avert production cloud VM interruptions. 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI '20), 1155–1170.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
All papers should be submitted electronically. All submitted manuscripts must be original work that is not under submission at another journal or under consideration for publication in another form, such as a monograph or chapter of a book. Authors of submitted papers are obligated not to submit their paper for publication elsewhere until an editorial decision is rendered on their submission. Further, authors of accepted papers are prohibited from publishing the results in other publications that appear before the paper is published in the Journal unless they receive approval for doing so from the Editor-In-Chief.
IJISAE open access articles are licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. This license lets the audience to give appropriate credit, provide a link to the license, and indicate if changes were made and if they remix, transform, or build upon the material, they must distribute contributions under the same license as the original.


