Self-Healing Cloud Data Pipelines Using AI-Based Failure Detection, Root Cause Analysis, and Automated Recovery

Authors

  • Pavan Kumar Kodati

Keywords:

automated recovery, cloud data pipelines, failure detection, root cause analysis, self-healing systems

Abstract

Enterprise cloud data pipelines fail far more often than organizations acknowledge, and the downstream cost of each failure extends well beyond the immediate job. Schema drift, delayed file arrivals, resource exhaustion, API instability, and orchestration defects each produce cascading effects across reporting layers, machine learning features, and regulatory deliverables. Current monitoring approaches detect failure states but stop short of diagnosis and remediation, leaving operational teams to manually cross-reference logs, lineage records, and infrastructure metrics before any recovery can begin. This article proposes a self-healing cloud pipeline framework that closes the gap between detection and resolution. The framework integrates observability telemetry aggregation, machine learning-based failure classification, retrieval-grounded large language model (LLM) root cause interpretation, and policy-driven automated recovery execution within a four-layer architecture. The layers, including observation, intelligence, decisioning, and recovery processes, manage incidents from the initial signal to the resolved state without requiring human intervention in approved scenarios. The approach targets measurable reductions in mean time to recovery (MTTR), lower manual intervention rates, and more consistent recovery behavior across large and diverse pipeline estates. The framework is designed for cloud-native orchestration environments and is assessed against rule-based and machine-learning-only baseline approaches across representative failure scenario categories.

Downloads

Download data is not yet available.

References

M. Zaharia et al., "Apache Spark: a unified engine for big data processing," Communications of the ACM, vol. 59, no. 11, pp. 56-65, 2016. [Online]. Available: https://dl.acm.org/doi/10.1145/2934664

J. Weng, J. H. Wang, J. Yang, and Y. Yang, "Root cause analysis of anomalies of multitier services in public clouds," in Proc. IEEE/ACM 25th International Symposium on Quality of Service (IWQoS), 2017, pp. 1-6. [Online]. Available: https://ieeexplore.ieee.org/document/7969155

R. Chaiken et al., "SCOPE: easy and efficient parallel processing of massive data sets," Proceedings of the VLDB Endowment, vol. 1, no. 2, pp. 1265-1276, 2008. [Online]. Available: https://doi.org/10.14778/1454159.1454166

F. Psallidas et al., "Provenance for interactive visualizations," in Proc. Workshop on Human-In-the-Loop Data Analytics, 2018, pp. 1-6. [Online]. Available: https://doi.org/10.1145/3209900.3209904

Y. Chen et al., "Outage prediction and diagnosis for cloud service systems," in Proc. The Web Conference 2019, 2019, pp. 2659-2665. [Online]. Available: https://doi.org/10.1145/3308558.3313501

P. Notaro et al., "A survey of AIOps methods for failure management," ACM Transactions on Intelligent Systems and Technology, vol. 13, no. 1, pp. 1-45, 2021. [Online]. Available: https://doi.org/10.1145/3483424

L. Ehrlinger and W. Woss, "Towards a definition of knowledge graphs," in Proc. SEMANTiCS 2016, 2016, pp. 1-4. [Online]. Available: https://ceur-ws.org/Vol-1695/paper4.pdf

M. Grottke and K. S. Trivedi, "Fighting bugs: remove, retry, replicate, and rejuvenate," IEEE Computer, vol. 40, no. 2, pp. 107-109, 2007. [Online]. Available: https://doi.org/10.1109/mc.2007.55

S. He et al., "A survey on automated log analysis for reliability engineering," ACM Computing Surveys, vol. 54, no. 6, pp. 1-37, 2021. [Online]. Available: https://doi.org/10.1145/3460345

J. Soldani and A. Brogi, "Anomaly detection and failure root cause analysis in (micro) service-based cloud applications," ACM Computing Surveys, vol. 55, no. 3, pp. 1-39, 2022. [Online]. Available: https://doi.org/10.1145/3501297

P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Proc. NeurIPS 2020, 2020, pp. 9459-9474. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf

A. Gulenko et al., "Detecting anomalous behavior of black-box services modeled with distance-based online clustering," in Proc. IEEE International Conference on Cloud Computing, 2018, pp. 912-915. [Online]. Available: https://doi.org/10.1109/cloud.2018.00134

M. Bosma et al., "Chain-of-thought prompting elicits reasoning in large language models," in Proc. NeurIPS 2022, 2022, pp. 24824-24837. [Online]. Available: https://doi.org/10.52202/068431-1800

S. He et al., "Towards automated log parsing for large-scale log data analysis," IEEE Transactions on Dependable and Secure Computing, vol. 15, no. 6, pp. 931-944, 2018. [Online]. Available: https://doi.org/10.1109/tdsc.2017.2762673

X. Zhou et al., "Latent error prediction and fault localization for microservice applications by learning from system trace logs," in Proc. 27th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE), 2019, pp. 683-694. [Online]. Available: https://dl.acm.org/doi/10.1145/3338906.3338961

G. Candea, S. Kawamoto, Y. Fujiki, G. Friedman, and A. Fox, "Microreboot: a technique for cheap recovery," in Proc. 6th Symposium on Operating Systems Design and Implementation (OSDI), 2004, pp. 31-44. [Online]. Available: https://www.usenix.org/legacy/event/osdi04/tech/full_papers/candea/candea.pdf

Downloads

Published

31.07.2026

How to Cite

Pavan Kumar Kodati. (2026). Self-Healing Cloud Data Pipelines Using AI-Based Failure Detection, Root Cause Analysis, and Automated Recovery. International Journal of Intelligent Systems and Applications in Engineering, 14(1s), 2128 –. Retrieved from https://ijisae.org/index.php/IJISAE/article/view/8488

Issue

Section

Research Article