Adaptive Autoscaling of Java Microservices Using Reinforcement Learning in Cloud-Native Environments
Keywords:
Microservice-based architectures, Horizontal Pod Autoscaler (HPA), reinforcement learning (RL), Java microservices, Markov Decision Process etc.Abstract
Microservice architectures have become the most common way to build enterprise scale cloud-native applications․ The default Kubernetes Horizontal Pod Autoscaler (HPA)‚ scales the application pods based on static CPU/Memory thresholds‚ which do not work well with variable‚ bursty workloads․ We build an adaptive reinforcement learning (RL)-based autoscaler for Java microservices that outputs optimal scaling policies by learning from runtime telemetry․ The DQN and Proximal Policy Optimisation (PPO) agents are trained on a 42-dimensional state space and a factored action space of 5^6 joint replica deltas via Markov Decision Process․ We use a hybrid trace consisting of the Alibaba Cluster Trace v2018 with synthetic bursty multipliers․ The multi-objective reward function includes penalties for violating the P95 latency‚ resource cost per pod‚ SLA violations‚ and under-utilization․ After training for over 2000 episodes‚ PPO attains a SLA compliance of 97․1%‚ a response time of 174 ms‚ and a reduction in cloud resource cost of 35․9% per hour compared to the static HPA․ Ablation studies confirm that all four reward terms are required for the Pareto-optimal policy․ This makes RL-based adaptive autoscaling a deployable alternative to the threshold-driven autoscaling schemes in production Kubernetes clusters․
Downloads
References
Burns, B., Grant, B., Oppenheimer, D., Brewer, E., and Wilkes, J. (2016). "Borg, Omega, and Kubernetes: Lessons learned from three container-management systems over a decade." ACM Queue, 14(1), 70–93.
Kubernetes Authors. (2023). "Horizontal Pod Autoscaling." Kubernetes Documentation. https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/.
Gandhi, R., Gupta, P., and Kumar, A. (2021). "Characterising autoscaling latency in Kubernetes under dynamic workloads." Proceedings of the 12th ACM Symposium on Cloud Computing (SoCC'21), pp. 301–313.
Hochreiter, S., and Schmidhuber, J. (1997). "Long short-term memory." Neural Computation, 9(8), 1735–1780.
Sutton, R. S., and Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press, Cambridge, MA.
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., … and Hassabis, D. (2015). "Human-level control through deep reinforcement learning." Nature, 518(7540), 529–533.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). "Proximal policy optimization algorithms." arXiv preprint arXiv:1707.06347.
Calheiros, R. N., Ranjan, R., Beloglazov, A., De Rose, C. A. F., and Buyya, R. (2011). "CloudSim: A toolkit for modeling and simulation of cloud computing environments and evaluation of resource provisioning algorithms." Software: Practice and Experience, 41(1), 23–50.
Shashank Daram, “A Review of Retrieval-Augmented Generation in Natural Language Processing: Architectures, Challenges, and Future Research Directions” , International Journal on Recent and Innovation Trends in Computing and Communication; vol 8 (12), 2020 :145-155.
Lim, H. C., Babu, S., Chase, J. S., and Parekh, S. (2009). "Automated control in cloud computing: Challenges and opportunities." Proceedings of the 1st Workshop on Automated Control for Datacenters and Clouds, pp. 13–18.
Cao, Z., Dong, S., Vemuri, S., and Du, D. H. C. (2020). "Characterizing, modeling, and predicting microservice dependencies for autoscaling in cloud-native applications." IEEE Transactions on Services Computing, 14(6), 1856–1869.
Toka, L., Dobreff, G., Fodor, B., and Sonkoly, B. (2021). "Adaptive machine learning-based autoscaling for heterogeneous workloads in Kubernetes environments." IEEE Transactions on Network and Service Management, 18(3), 3306–3320.
Nguyen, T., Yeow, J., Phan, T., and Nguyen, M. (2022). "Container resource right-sizing using Bayesian optimisation for microservice deployments." Journal of Systems and Software, 184, 111128.
Mao, H., Alizadeh, M., Menache, I., and Kandula, S. (2016). "Resource management with deep reinforcement learning." Proceedings of the 15th ACM Workshop on Hot Topics in Networks (HotNets'16), pp. 50–56.
Peng, Z., Lin, H., and Zhang, W. (2019). "Virtual machine consolidation using multi-agent deep reinforcement learning in private cloud environments." IEEE Access, 7, 83927–83940.
Rzadca, K., et al. (2020). "Autopilot: Workload autoscaling at Google." Proceedings of the European Conference on Computer Systems (EuroSys'20), pp. 1–16.
Shashank Lohani, The Learning Management System as a Product: A Computer Science Perspective on Higher Education Strategy, International Journal of Communication Networks and Information Security; vol 11 (2), 2019 page no 353-363.
Chen, W., Shi, J., Li, B., and Li, Z. (2022). "An actor-critic reinforcement learning approach for Kubernetes pod autoscaling in multi-tenant serverless environments." Future Generation Computer Systems, 134, 97–111.
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N. (2016). "Dueling network architectures for deep reinforcement learning." Proceedings of ICML 2016, Vol. 48, pp. 1995–2003.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
All papers should be submitted electronically. All submitted manuscripts must be original work that is not under submission at another journal or under consideration for publication in another form, such as a monograph or chapter of a book. Authors of submitted papers are obligated not to submit their paper for publication elsewhere until an editorial decision is rendered on their submission. Further, authors of accepted papers are prohibited from publishing the results in other publications that appear before the paper is published in the Journal unless they receive approval for doing so from the Editor-In-Chief.
IJISAE open access articles are licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. This license lets the audience to give appropriate credit, provide a link to the license, and indicate if changes were made and if they remix, transform, or build upon the material, they must distribute contributions under the same license as the original.


