AIOps-Driven Anomaly Detection Framework for Enterprise Infrastructure Reliability Using Time-Series Forecasting and Intelligent Alert Correlation

Authors

  • Muhammed Yunas Muhammed Yunas Chirayath Meerankunju

Keywords:

AIOps, anomaly detection, time-series forecasting, observability, site reliability engineering, Datadog, Moogsoft, Isolation Forest, Prophet, SARIMA, alert correlation, financial infrastructure, DBSCAN, MTTD, MTTR, SLO engineering.

Abstract

Static-threshold monitoring decays predictably as financial infrastructure scales -- a fundamental structural failure that no amount of tuning resolves [9]. This paper addresses this gap through a novel combination: seasonality-aware time-series forecasting integrated directly with topology-aware alert correlation in a single production-calibrated pipeline. Unlike prior AIOps systems that treat forecasting and correlation as independent modules or rely on proprietary vendor platforms without generalizable design specifications [1][3], the contribution here is their joint calibration to the distinctive periodicity of payment workloads and the deployment of both within a vendor-neutral, reproducible architecture. The novelty lies in three integrated design decisions absent from existing systems: (1) Prophet-based seasonal decomposition [6] pre-conditioned on payment calendar events (end-of-month settlement, intraday volume peaks) as explicit changepoint priors rather than post-hoc corrections; (2) Isolation Forest multivariate scoring [8] calibrated on a topology-derived metric neighborhood rather than independent feature sets; and (3) APM-trace-derived service graphs used as the spatial metric for DBSCAN situation clustering [17], directly coupling the anomaly detection and correlation layers. The system was validated on an enterprise financial platform processing over 40,000 transactions per second, using an OpenTelemetry-compatible collection layer, a Kafka-based ingestion pipeline, and a commercial event correlation platform (Moogsoft-compatible architecture) across more than 12,000 monitored hosts. Over a 90-day production evaluation against 180 ground-truth incidents, the system achieved precision of 0.91 and recall of 0.94 (F1=0.92, 95% CI: [0.90, 0.94]), reduced total daily alert volume by 75%, increased actionable alert volume by 77%, and cut the false positive rate from 87.2% to 9.4% [1][8]. MTTD improved by 82% and MTTR by 74% across incident categories. We document implementation decisions, failure modes, and the organizational dynamics that shaped these outcomes.

Downloads

Download data is not yet available.

References

P. Notaro, J. Cardoso, and M. Gerndt, 'A Survey of AIOps Methods for Failure Management,' ACM Trans. Intell. Syst. Technol., vol. 12, no. 6, Art. no. 72, Nov. 2021, doi:10.1145/3465485.

M. Chen, X. Zheng, J. Lloyd, M. I. Jordan, and E. Brewer, 'Failure Diagnosis Using Decision Trees,' in Proc. IEEE Int. Conf. Autonomic Computing, 2004, pp. 36-43.

S. Dang, Q. Lin, J.-G. Lou, H. Zhang, D. Qiao, and W. Xu, 'AIOps: Real-World Challenges and Research Innovations,' in Proc. ICSE Companion, 2019, pp. 4-5.

Datadog, Inc., 'Datadog Infrastructure Monitoring Documentation,' 2023. [Online]. Available: https://docs.datadoghq.com/infrastructure/

Moogsoft, Inc., 'Moogsoft AIOps Platform Documentation,' 2023. [Online]. Available: https://docs.moogsoft.com/

S. J. Taylor and B. Letham, 'Forecasting at Scale,' Am. Statistician, vol. 72, no. 1, pp. 37-45, 2018, doi:10.1080/00031305.2017.1380080.

G. E. P. Box, G. M. Jenkins, G. C. Reinsel, and G. M. Ljung, Time Series Analysis: Forecasting and Control, 5th ed. Hoboken, NJ: Wiley, 2015.

F. T. Liu, K. M. Ting, and Z. H. Zhou, 'Isolation Forest,' in Proc. IEEE ICDM, 2008, pp. 413-422, doi:10.1109/ICDM.2008.17.

C. H. Lim, D. Yu, P. Krishnamurthy, Y. Kim, C. Liu, and R. Alrawis, 'Poster: Towards Continuous Threshold Optimization for Microservice Monitoring,' in Proc. ASE, 2020.

V. Nair, T. Menzies, N. Siegmund, and S. Apel, 'Faster Discovery of Faster System Configurations with Spectral Learning,' Autom. Softw. Eng., vol. 25, no. 2, pp. 247-277, 2018.

G. Soldani, D. A. Tamburri, and W.-J. Van Den Heuvel, 'The Pains and Gains of Microservices: A Systematic Grey Literature Review,' J. Syst. Softw., vol. 146, pp. 215-232, Dec. 2018.

P. Brockwell and R. Davis, Introduction to Time Series and Forecasting, 3rd ed. New York, NY: Springer, 2016.

S. Hochreiter and J. Schmidhuber, 'Long Short-Term Memory,' Neural Comput., vol. 9, no. 8, pp. 1735-1780, 1997.

S. Mariani, G. Omicini, and G. Agha, 'Tuple-Based Coordination in a Distributed Setting,' in Proc. 21st Euromicro Int. Conf. Parallel, Distributed, Network-Based Processing, 2013.

R. Kuppusamy, 'Tracking Microservice Dependencies,' Netflix Technology Blog, 2022. [Online]. Available: https://netflixtechblog.com/

B. Mishchenko, 'Eliminating Toil with Fully Automated Deployments,' LinkedIn Engineering Blog, 2023. [Online]. Available: https://engineering.linkedin.com/

M. Ester, H.-P. Kriegel, J. Sander, and X. Xu, 'A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise,' in Proc. KDD, 1996, pp. 226-231.

B. Beyer, C. Jones, J. Petoff, and N. R. Murphy, Site Reliability Engineering: How Google Runs Production Systems. Sebastopol, CA: O'Reilly Media, 2016.

B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, and S. Thorne, The Site Reliability Workbook. Sebastopol, CA: O'Reilly Media, 2018.

J. Dean and L. A. Barroso, 'The Tail at Scale,' Commun. ACM, vol. 56, no. 2, pp. 74-80, Feb. 2013.

W. Xu, L. Huang, A. Fox, D. Patterson, and M. I. Jordan, 'Detecting Large-Scale System Problems by Mining Console Logs,' in Proc. ACM SOSP, 2009, pp. 117-132.

Q. Lin, H. Zhang, J.-G. Lou, Y. Zhang, and X. Chen, 'Log Clustering Based Problem Identification for Online Service Systems,' in Proc. ICSE-C, 2016.

H. Mi, H. Wang, Y. Zhou, M. R. Lyu, and H. Cai, 'Toward Fine-Grained, Unsupervised, Scalable Performance Diagnosis for Production Cloud Computing Systems,' IEEE Trans. Parallel Distrib. Syst., vol. 24, no. 6, pp. 1245-1255, Jun. 2013.

M. Kleppmann, Designing Data-Intensive Applications. Sebastopol, CA: O'Reilly Media, 2017.

OpenTelemetry Authors, 'OpenTelemetry Specification v1.x,' CNCF, 2023. [Online]. Available: https://opentelemetry.io/docs/

Apache Software Foundation, 'Apache Kafka Documentation,' 2023. [Online]. Available: https://kafka.apache.org/documentation/

Apache Software Foundation, 'Apache Flink Documentation,' 2023. [Online]. Available: https://nightlies.apache.org/flink/flink-docs-stable/

T. Chen, X. Zhang, and H. Zhang, 'Outage Prediction and Diagnosis for Cloud Service Systems,' in Proc. WWW, 2019, pp. 2659-2665.

PCI Security Standards Council, 'PCI DSS v4.0,' 2022. [Online]. Available: https://www.pcisecuritystandards.org/

R. Syer, Z. Shang, H. Jiang, and A. E. Hassan, 'Continuous Validation of Performance Tests by Statistical Analysis of Performance Evolution,' J. Syst. Softw., vol. 103, pp. 224-237, May 2015.

X. Wang, L. Wu, M. Yin, Y. Wen, and Y. Li, 'Groot: An Event-Graph-Based Approach for Root Cause Analysis in Industrial Settings,' in Proc. ASE, 2021, pp. 1219-1230.

Amazon Web Services, 'AWS Observability Best Practices,' 2023. [Online]. Available: https://aws-observability.github.io/observability-best-practices/

R. P. Adams and D. J. C. MacKay, 'Bayesian Online Changepoint Detection,' arXiv:0710.3742, 2007. [Online]. Available: https://arxiv.org/abs/0710.3742

Downloads

Published

30.05.2023

How to Cite

Muhammed Yunas Muhammed Yunas Chirayath Meerankunju. (2023). AIOps-Driven Anomaly Detection Framework for Enterprise Infrastructure Reliability Using Time-Series Forecasting and Intelligent Alert Correlation. International Journal of Intelligent Systems and Applications in Engineering, 11(6s), 1014–1030. Retrieved from https://ijisae.org/index.php/IJISAE/article/view/8461

Issue

Section

Research Article