Cloud-Scale Data Engineering Using PySpark and Delta Lake: Architecture, Performance, and Governance
Keywords:
Cloud data engineering, Apache PySpark, Delta Lake, lakehouse architecture, distributed data processing, cloud computing, data governance, enterprise intelligence, query performance optimisation, storage reliability, real-time analytics, scalable data architectureAbstract
Cloud data engineering has become essential for many modern enterprises as organisations are increasingly adopting cloud solutions to handle their large-scale data and execute analytical processes. This study analysed scalable cloud data engineering using PySpark and Delta Lake by exploring the implication on architectural scalability, computational performance, storage reliability, governance, and enterprise intelligence. The paper employed a qualitative research design using an inductive approach by analysing secondary data from various sources, including journals, articles, conference proceedings, and reports. Articles published between 2020 and 2026 provided the basis for discussing the impact of deploying PySpark and Delta Lake on cloud data engineering. The study found that PySpark improved distributed data processing using in-memory computations while enabling faster analytical query execution, thus enhancing parallel processing capabilities with increased data volume. Moreover, the study found that Delta Lake improved storage reliability by implementing ACID transactions, journaling, optimistic concurrency, and versioning. The researcher also learned that performance optimisation through adaptive query execution, partitioning, caching, file compaction and optimised MERGE operations improved query processing and computational performance. Finally, it was concluded that governance capabilities such as metadata management, encryption, access controls, and regulatory compliance increased organisational confidence in the use of cloud storage, while also ensuring data security. The study ultimately concludes that the combined application of PySpark and Delta Lake can improve a lakehouse platform that executes enterprise-level data engineering, analytics, governance, and decision-support activities while responding to the growing complexity of modern cloud data systems.
Downloads
References
Armbrust, M., Das, T., Sun, L., Yavuz, B., Zhu, S., Murthy, M., Torres, J., Van Hovell, H., Ionescu, A., Łuszczak, A. and Świtakowski, M., 2020. Delta lake: high-performance ACID table storage over cloud object stores. Proceedings of the VLDB Endowment, 13(12), pp.3411-3424. Available at https://www.eecs.umich.edu/courses/cse584/static_files/papers/p975-armbrust.pdf
Armbrust, M., Das, T., Sun, L., Yavuz, B., Zhu, S., Murthy, M., Torres, J., Van Hovell, H., Ionescu, A., Łuszczak, A. and Świtakowski, M., 2020. Delta lake: high-performance ACID table storage over cloud object stores. Proceedings of the VLDB Endowment, 13(12), pp.3411-3424. Available at https://www.eecs.umich.edu/courses/cse584/static_files/papers/p975-armbrust.pdf
Arul, K., 2022. Data Engineering Challenges in Multi-cloud Environments: Strategies for Efficient Big Data Integration and Analytics. International Journal of Scientific Research and Management (IJSRM), 10(06). Available at https://www.academia.edu/download/123451259/Data_Engineering_Challenges_in_Multi_2_.pdf
Belur, J., Tompson, L., Thornton, A. and Simon, M., 2021. Interrater reliability in systematic review methodology: exploring variation in coder decision-making. Sociological methods & research, 50(2), pp.837-865. Available at https://journals.sagepub.com/doi/abs/10.1177/0049124118799372
Chai, Y., Chai, Y., Wang, X., Wei, H., Bao, N. and Liang, Y., 2019, April. LDC: a lower-level driven compaction method to optimize SSD-oriented key-value stores. In 2019 IEEE 35th International Conference on Data Engineering (ICDE) (pp. 722-733). IEEE. https://arxiv.org/abs/2004.03054
Chinta, S., 2022. Integrating artificial intelligence with cloud business intelligence: Enhancing predictive analytics and data visualization. Iconic Research And Engineering Journals, 5(9). Available at https://www.academia.edu/download/119711197/Integrating_Artificial_Intelligence_sweth.pdf
Chowdhury, R.H., 2021. Cloud-based data engineering for scalable business analytics solutions: designing scalable cloud architectures to enhance the efficiency of big data analytics in enterprise settings. Journal of Technological Science & Engineering (JTSE), 2(1), pp.21-33. Available at https://www.rsepress.org/index.php/jtse/article/view/93
Chui, M., Hall, B., Mayhew, H., Singla, A. and Sukharevsky, A. (2022) The state of AI in 2022—and a half decade in review. McKinsey & Company, 6 December. Available at: https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-in-2022-and-a-half-decade-in-review
Databricks (2021) The foundation of your lakehouse starts with Delta Lake. Databricks Blog, 1 December. Available at: https://www.databricks.com/blog/2021/12/01/the-foundation-of-your-lakehouse-starts-with-delta-lake.html
Davidson, E., Edwards, R., Jamieson, L. and Weller, S., 2019. Big data, qualitative style: a breadth-and-depth method for working with large amounts of secondary qualitative data. Quality & quantity, 53(1), pp.363-376. Available at https://link.springer.com/article/10.1007/s11135-018-0757-y
Kaul, D., 2019. Optimizing resource allocation in multi-cloud environments with artificial intelligence: Balancing cost, performance, and security. JICET, 4, pp.1-25. Available at https://www.academia.edu/download/131654901/OptimizingResourceAllocationinMulti_CloudEnvironments.pdf
Kelsoe, C., 2019. Estimating recreational value of water quality in Mississippi lakes when water quality data are scarce. Available at https://scholarsjunction.msstate.edu/td/1930/
Mohna, H.A., Barua, T., Mohiuddin, M. and Rahman, M.M., 2022. AI-ready data engineering pipelines: a review of medallion architecture and cloud-based integration models. American Journal of Scholarly Research and Innovation, 1(01), pp.319-350. Available at https://researchinnovationjournal.com/index.php/AJSRI/article/view/51
Mohna, H.A., Barua, T., Mohiuddin, M. and Rahman, M.M., 2022. AI-ready data engineering pipelines: a review of medallion architecture and cloud-based integration models. American Journal of Scholarly Research and Innovation, 1(01), pp.319-350. Available at https://researchinnovationjournal.com/index.php/AJSRI/article/view/51
Mudusu, S.K., 2022. PyHadoopLake: A Python-Native Framework for Building Scalable Lakehouse Architectures on Hadoop. International Journal of Research Publications in Engineering, Technology and Management (IJRPETM), 5(5), pp.7449-7452. Available at https://www.ijrpetm.com/index.php/IJRPETM/article/view/473
Nagabhyru, K.C., 2022. Bridging Traditional ETL Pipelines with AI Enhanced Data Workflows: Foundations of Intelligent Automation in Data Engineering. Available at SSRN 5505199. Available at https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5505199
Ross, R., Pillitteri, V., Graubart, R., Bodeau, D. and McQuaid, R., 2019. Developing cyber resilient systems: a systems security engineering approach (No. NIST Special Publication (SP) 800-160 Vol. 2 (Draft)). National Institute of Standards and Technology. Available at https://csrc.nist.gov/CSRC/media/Publications/sp/800-160/vol-2/draft/documents/sp800-160-vol2-draft-fpd.pdf
Shanbhag, A., Madden, S. and Yu, X., 2020, June. A study of the fundamental performance characteristics of GPUs and CPUs for database analytics. In Proceedings of the 2020 ACM SIGMOD international conference on Management of data (pp. 1617-1632). Available at https://dl.acm.org/doi/abs/10.1145/3318464.3380595
Tufă, A., 2019. Self-organizing data layouts for databricks delta. CWI, 20, p.29. Available at https://homepages.cwi.nl/~boncz/msc/2019-AdrianaTufa.pdf
Sibeoni, J., Verneuil, L., Manolios, E. and Révah-Levy, A., 2020. A specific method for qualitative medical research: the IPSE (Inductive Process to analyze the Structure of lived Experience) approach. BMC Medical Research Methodology, 20(1), p.216. Available at https://link.springer.com/article/10.1186/s12874-020-01099-4
Suri-Payer, F., Burke, M., Wang, Z., Zhang, Y., Alvisi, L. and Crooks, N., 2021, October. Basil: Breaking up BFT with ACID (transactions). In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles (pp. 1-17). Available at https://dl.acm.org/doi/abs/10.1145/3477132.3483552
Hu, Y., Zhu, Z., Neal, I., Kwon, Y., Cheng, T., Chidambaram, V. and Witchel, E., 2019. TxFS: Leveraging file-system crash consistency to provide ACID transactions. ACM Transactions on Storage (TOS), 15(2), pp.1-20. Available at https://dl.acm.org/doi/abs/10.1145/3318159
Synergy Research Group (2021) 2020 – The year that cloud service revenues finally dwarfed enterprise spending on data centers. Synergy Research Group, 18 March. Available at: https://www.srgresearch.com/articles/2020-the-year-that-cloud-service-revenues-finally-dwarfed-enterprise-spending-on-data-centers
Zakarya, M., Gillam, L., Ali, H., Rahman, I.U., Salah, K., Khan, R., Rana, O. and Buyya, R., 2020. epcAware: A game-based, energy, performance and cost-efficient resource management technique for multi-access edge computing. IEEE transactions on services computing, 15(3), pp.1634-1648. Available at https://ieeexplore.ieee.org/abstract/document/9127143/
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
All papers should be submitted electronically. All submitted manuscripts must be original work that is not under submission at another journal or under consideration for publication in another form, such as a monograph or chapter of a book. Authors of submitted papers are obligated not to submit their paper for publication elsewhere until an editorial decision is rendered on their submission. Further, authors of accepted papers are prohibited from publishing the results in other publications that appear before the paper is published in the Journal unless they receive approval for doing so from the Editor-In-Chief.
IJISAE open access articles are licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. This license lets the audience to give appropriate credit, provide a link to the license, and indicate if changes were made and if they remix, transform, or build upon the material, they must distribute contributions under the same license as the original.


