Evaluating PHI Leakage Risks and Governance Mechanisms in Healthcare Generative AI Pipelines

Authors

  • Vinod Rufus Motani

Keywords:

Generative Artificial Intelligence (GenAI), Protected Health Information (PHI), Healthcare AI, Data Privacy, PHI Leakage, AI Governance, HIPAA Compliance, Large Language Models (LLMs), Healthcare Data Security, Privacy-Preserving AI.

Abstract

There has been an increased use of Generative Artificial Intelligence (GenAI) in healthcare to assist in documentation, diagnosis, communication with patients, and decision-making in administration. However, the increase in the use of this technology has resulted in some major risks with respect to PHI leakage through AI pipelines. In this research, PHI leakage risks and governance strategies in healthcare generative AI pipelines have been investigated using a qualitative research approach. Secondary data were gathered from peer-reviewed journal articles, healthcare regulations, policy papers, and AI governance literature. Inductive research approach and thematic analysis were applied to recognize major privacy risks and governance strategies across various stages of the AI pipeline. As a result of this analysis, the risks of PHI leakage at stages of data collection, preprocessing, model training, fine-tuning, and inference have been identified. Major vulnerabilities include insecure cloud storage, weak access control, prompt-based PHI leakage, membership inference attacks, data memorisation, and risks of retrieval-augmented generation. Moreover, organizational procedures such as continuous monitoring of models, staff education, and standardized governance framework of AI contribute to privacy compliance and security of AI implementation. In conclusion, privacy protection of PHI requires governance measures throughout the whole generative AI life cycle instead of only at one particular phase of development. Combining technical controls and regulatory and governance measures can lead to lowering privacy threats and fostering secure and trustworthy adoption of generative AI in healthcare organizations.

Downloads

Download data is not yet available.

References

Agrawal, N., Binns, R., Van Kleek, M., Laine, K., & Shadbolt, N. (2021, May). Exploring design and governance challenges in the development of privacy-preserving computation. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (pp. 1-13). https://dl.acm.org/doi/abs/10.1145/3411764.3445677

Ahmed, Z., Mohamed, K., Zeeshan, S., & Dong, X. (2020). Artificial intelligence with multi-functional machine learning platform development for better healthcare and precision medicine. Database, 2020, baaa010. https://academic.oup.com/database/article-abstract/doi/10.1093/database/baaa010/5809229

Azizi, Z., Zheng, C., Mosquera, L., Pilote, L., El Emam, K., & GOING-FWD Collaborators. (2021). Can synthetic data be a proxy for real clinical trial data? A validation study. BMJ open, 11(4), e043497. https://bmjopen.bmj.com/content/11/4/e043497.abstract

Bosch, J., Olsson, H. H., & Crnkovic, I. (2021). Engineering ai systems: A research agenda. Artificial intelligence paradigms for smart cyber-physical systems, 1-19. https://www.igi-global.com/chapter/engineering-ai-systems/266130

Dokuchaev, V. A., Maklachkova, V. V., & Statev, V. Y. (2020). Classification of personal data security threats in information systems. T-Comm-Телекоммуникации и Транспорт, 14(1), 56-60. https://cyberleninka.ru/article/n/classification-of-personal-data-security-threats-in-information-systems

Eleanor, H. (2021). Modernizing data security: Best practices for compliance with US and international privacy regulations. International Journal of Trend in Scientific Research and Development, 5(4), 1881-1894. http://eprints.umsida.ac.id/16095/

Guidance, W. H. O. (2021). Ethics and governance of artificial intelligence for health. World Health Organization, 1. https://journals.sagepub.com/doi/abs/10.3233/SHTI200542

Hu, Y., Kuang, W., Qin, Z., Li, K., Zhang, J., Gao, Y., ... & Li, K. (2021). Artificial intelligence security: Threats and countermeasures. ACM Computing Surveys (CSUR), 55(1), 1-36. https://dl.acm.org/doi/abs/10.1145/3487890

Khan, Z. F., & Alotaibi, S. R. (2020). Applications of artificial intelligence and big data analytics in m‐health: A healthcare system perspective. Journal of healthcare engineering, 2020(1), 8894694. https://onlinelibrary.wiley.com/doi/abs/10.1155/2020/8894694

Labra, O., Castro, C., Wright, R., & Chamblas, I. (2020). Thematic analysis in social work: A case study. Global social work-cutting edge issues and critical reflections, 10(6), 1-20. https://books.google.com/books?hl=en&lr=&id=UBT8DwAAQBAJ&oi=fnd&pg=PA183&dq=Thematic+analysis+technique+was+adopted+to+analyse+the+gathered+evidence&ots=ASmUBBV2HI&sig=dBhCq9-yxyzgtC3nQrYKLv6gs1Y

Lee, T., B., &Trott, S., 2023. Large language models, explained with a minimum of math and jargon. https://www.understandingai.org/p/large-language-models-explained-with

Lehmkuhl, R., Mishra, P., Srinivasan, A., & Popa, R. A. (2021). Muse: Secure inference resilient to malicious clients. In 30th USENIX Security Symposium (USENIX Security 21) (pp. 2201-2218). https://www.usenix.org/conference/usenixsecurity21/presentation/lehmkuhl

Personal Data Protection Commission. (2020). Model AI Governance Framework (2020). https://openresearch-repository.anu.edu.au/bitstreams/00df4e3b-8176-4497-8384-a55e72c20f56/download

Prado, M. D., Su, J., Saeed, R., Keller, L., Vallez, N., Anderson, A., ... & Pazos, N. (2020). Bonseyes ai pipeline—bringing ai to you: End-to-end integration of data, algorithms, and deployment tools. ACM Transactions on Internet of Things, 1(4), 1-25. https://dl.acm.org/doi/abs/10.1145/3403572

Ramponi, M., & O'Connor, R., 2023. The Full Story of Large Language Models and RLHF. https://www.assemblyai.com/blog/the-full-story-of-large-language-models-and-rlhf

Shah, S. M., & Khan, R. A. (2020). Secondary use of electronic health record: Opportunities and challenges. IEEE access, 8, 136947-136965. https://ieeexplore.ieee.org/abstract/document/9146114/

Solanki, P., Grundy, J., & Hussain, W. (2023). Operationalising ethics in artificial intelligence for healthcare: A framework for AI developers. AI and Ethics, 3(1), 223–240. https://doi.org/10.1007/s43681-022-00195-z

Topaloglu, M. Y., Morrell, E. M., Rajendran, S., & Topaloglu, U. (2021). In the pursuit of privacy: the promises and predicaments of federated learning in healthcare. Frontiers in Artificial Intelligence, 4, 746497. https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2021.746497/full

Truby, J. (2020). Governing artificial intelligence to benefit the UN sustainable development goals. Sustainable Development, 28(4), 946-959. https://onlinelibrary.wiley.com/doi/abs/10.1002/sd.2048

Wannamaker, K. (2021). Situated Self-Tracking: Ideating, Designing, and Deploying Dedicated User-driven Personal Informatics Systems. https://ucalgary.scholaris.ca/items/77a91ad5-e560-4ae2-a5a6-b18388c64723

Xiang, D., & Cai, W. (2021). Privacy protection and secondary use of health data: strategies and methods. BioMed Research International, 2021(1), 6967166. https://onlinelibrary.wiley.com/doi/abs/10.1155/2021/6967166

Xu, W., & Zammit, K. (2020). Applying thematic analysis to education: A hybrid approach to interpreting data in practitioner research. International journal of qualitative methods, 19, 1609406920918810. https://journals.sagepub.com/doi/abs/10.1177/1609406920918810

Downloads

Published

30.04.2023

How to Cite

Vinod Rufus Motani. (2023). Evaluating PHI Leakage Risks and Governance Mechanisms in Healthcare Generative AI Pipelines. International Journal of Intelligent Systems and Applications in Engineering, 11(5s), 692–698. Retrieved from https://ijisae.org/index.php/IJISAE/article/view/8553