Machine Learning-Based Predictive Fault Detection and Resilience Optimization in Large-Scale Distributed Systems

Authors

  • Ishu Anand Jaiswal

Keywords:

Machine Learning, Predictive Fault Detection, Distributed Systems, Fault Tolerance, Resilience Optimization.

Abstract

The rapid expansion of cloud computing, high-performance computing, large-scale data centers, distributed applications, and service-oriented computing infrastructures has significantly increased the complexity of modern distributed systems. Such environments consist of large numbers of interconnected computing nodes, virtual machines, storage devices, network resources, and software services that operate collaboratively to deliver continuous and scalable computing services. However, the increasing scale, heterogeneity, workload variability, and dynamic behavior of distributed infrastructures also increase their exposure to hardware faults, software failures, node crashes, resource exhaustion, communication disruptions, and performance degradation. Traditional fault-tolerance mechanisms generally react to failures after they occur and therefore may result in service interruption, resource wastage, recovery overhead, and violations of service-level requirements.

Downloads

Download data is not yet available.

References

Salami, H., Saadatfar, H., Rahmani Fard, F., Shekofteh, S. K., & Deldari, H. (2010). Improving cluster computing performance based on job futurity prediction. 2010 3rd International Conference on Advanced Computer Theory and Engineering (ICACTE), Vol. 6, V6-303–V6-307. https://doi.org/10.1109/ICACTE.2010.5579820

Saadatfar, H., Fadishei, H., & Deldari, H. (2012). Predicting job failures in AuverGrid based on workload log analysis. New Generation Computing, 30(1), 73–94. https://doi.org/10.1007/s00354-012-0105-z

Chen, X., Lu, C.-D., & Pattabiraman, K. (2014). Failure analysis of jobs in compute clouds: A Google cluster case study. IEEE 25th International Symposium on Software Reliability Engineering (ISSRE), 167–177. https://doi.org/10.1109/ISSRE.2014.34

Chen, X., Lu, C.-D., & Pattabiraman, K. (2014). Failure prediction of jobs in compute clouds: A Google cluster case study. IEEE 25th International Symposium on Software Reliability Engineering Workshops (ISSREW), 341–346. https://doi.org/10.1109/ISSREW.2014.105

Soualhia, M., Khomh, F., & Tahar, S. (2015). Predicting scheduling failures in the cloud: A case study with Google clusters and Hadoop on Amazon EMR. 2015 IEEE 17th International Conference on High Performance Computing and Communications, 58–65. https://doi.org/10.1109/HPCC-CSS-ICESS.2015.170

Jaiswal, I. A. (2020). Machine learning-based predictive fault detection and resilience optimization in large-scale distributed systems. International Journal of Intelligent Systems and Applications in Engineering.

Zheng, W., Wang, Z., Huang, H., Meng, L., & Qiu, X. (2016). EHMM-CT: An online method for failure prediction in cloud computing systems. KSII Transactions on Internet and Information Systems, 10(9), 4087–4107. https://doi.org/10.3837/tiis.2016.09.004

Islam, T., & Manivannan, D. (2017). Predicting application failure in cloud: A machine learning approach. 2017 IEEE 1st International Conference on Cognitive Computing (ICCC), 24–31. https://doi.org/10.1109/IEEE.ICCC.2017.11

Pitakrat, T., Okanović, D., van Hoorn, A., & Grunske, L. (2018). Hora: Architecture-aware online failure prediction. Journal of Systems and Software, 137, 669–685. https://doi.org/10.1016/j.jss.2017.02.041

Lin, Q., Hsieh, K., Dang, Y., Zhang, H., Sui, K., Xu, Y., Lou, J.-G., Li, C., Wu, Y., Yao, R., Chintalapati, M., & Zhang, D. (2018). Predicting node failure in cloud service systems. Proceedings of the 26th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 480–490. https://doi.org/10.1145/3236024.3236060

Mariani, L., Pezzè, M., Riganelli, O., & Xin, R. (2020). Predicting failures in multi-tier distributed systems. Journal of Systems and Software, 161, 110464. https://doi.org/10.1016/j.jss.2019.110464

Gummadi, V. P. K. (2020). API design and implementation: RAML and OpenAPI specification. Journal of Electrical Systems, 16(4). https://doi.org/10.52783/jes.9329

Lu, S., Luo, B., Patel, T., Yao, Y., Tiwari, D., & Shi, W. (2020). Making disk failure predictions SMARTer! 18th USENIX Conference on File and Storage Technologies (FAST '20), 151–167.

Ahmad, W., Khan, S. A., Kim, C. H., & Kim, J.-M. (2020). Feature selection for improving failure detection in hard disk drives using a genetic algorithm and significance scores. Applied Sciences, 10(9), 3200. https://doi.org/10.3390/app10093200

Li, Z., Cheng, Q., Hsieh, K., Dang, Y., Huang, P., Singh, P., Yang, X., Lin, Q., Wu, Y., Levy, S., & Chintalapati, M. (2020). Gandalf: An intelligent, end-to-end analytics service for safe deployment in large-scale cloud infrastructure. 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI '20), 389–402.

Levy, S., Yao, R., Wu, Y., Dang, Y., Huang, P., Mu, Z., Zhao, P., Ramani, T., Govindaraju, N., Li, X., Lin, Q., Shafriri, G. L., & Chintalapati, M. (2020). Predictive and adaptive failure mitigation to avert production cloud VM interruptions. 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI '20), 1155–1170.

Downloads

Published

30.12.2020

How to Cite

Ishu Anand Jaiswal. (2020). Machine Learning-Based Predictive Fault Detection and Resilience Optimization in Large-Scale Distributed Systems. International Journal of Intelligent Systems and Applications in Engineering, 8(4), 453–462. Retrieved from https://ijisae.org/index.php/IJISAE/article/view/8529

Issue

Section

Research Article