Understanding Artificial Intelligence Ethics and Safety: Foundations, Challenges, and Governance

Authors

  • Muhammad Sami Intizar Muhammad Nawaz Sharif University of Engineering & Technology image/svg+xml
  • Aqsa Siddique Regional Institute of Allied Health Sciences
  • Maria Sikandar Hamdard University image/svg+xml
  • Madiha Sikandar University of Central Punjab image/svg+xml
  • Maria Farooq Regional Institute of Allied Health Sciences

DOI:

https://doi.org/10.63158/IJAIS.v3i1.61

Keywords:

AI Ethics, AI Safety, Algorithmic Fairness, Value Alignment, AI Governance, Trustworthy AI

Abstract

The deployment of artificial intelligence (AI) in safety-critical domains has made ethics and safety priorities. This paper examines AI ethics and safety by synthesizing technical, philosophical, and governance perspectives. It analyzes ethical principles, including fairness, transparency, privacy, accountability, and safety, while examining tensions in their implementation. Technical approaches, including robustness mechanisms, adversarial training, and interpretability methods, are considered alongside frameworks for value alignment and corrigibility. The discussion highlights the challenges of translating values into operational requirements, when competing objectives and contextual differences complicate implementation. The paper also assesses the governance landscape, comparing regulatory approaches across jurisdictions and evaluating binding and non-binding instruments for AI development. It argues that governance requires integrating technical safeguards, ethical reflection, and institutional accountability throughout the AI lifecycle. Moving beyond principle-based declarations, the paper emphasizes auditable and enforceable practices that connect ethical commitments with implementation and oversight. This perspective supports responsible innovation while addressing the risks associated with safety-critical AI deployment.

References

[1] D. Mbiazi, M. Bhange, M. Babaei, I. Sheth, P. Kenfack, and S. E. Kahou, "Survey on AI ethics: A socio-technical perspective," Comput. Intell., vol. 41, no. 6, p. e70149, 2025.

[2] A. Jobin, M. Ienca, and E. Vayena, "The global landscape of AI ethics guidelines," Nat. Mach. Intell., vol. 1, no. 9, pp. 389–399, 2019. doi: 10.1038/s42256-019-0088-2

[3] M. Ienca and E. Vayena, "AI ethics guidelines: European and global perspectives," in Towards Regulation of AI Systems, Cham, Switzerland: Springer, 2020, pp. 38–60.

[4] L. Floridi and J. Cowls, "A unified framework of five principles for AI in society," in Machine Learning and the City: Applications in Architecture and Urban Design, Hoboken, NJ, USA: Wiley, 2022, pp. 535–545. doi: 10.1002/9781119815075.ch45

[5] J. Whittlestone, R. Nyrup, A. Alexandrova, and S. Cave, "The role and limits of principles in AI ethics: Towards a focus on tensions," in Proc. 2019 AAAI/ACM Conf. AI, Ethics, Soc. (AIES), 2019, pp. 195–200. doi: 10.1145/3306618.3314289

[6] D. Danks, "Governance via explainability," in The Oxford Handbook of AI Governance, Oxford, U.K.: Oxford University Press, 2022, pp. 183–197.

[7] B. Goodman and S. Flaxman, "European Union regulations on algorithmic decision-making and a 'right to explanation'," AI Mag., vol. 38, no. 3, pp. 50–57, 2017. doi: 10.1609/aimag.v38i3.2741

[8] B. Hedden, "On statistical criteria of algorithmic fairness," Phil. & Pub. Aff., vol. 49, no. 2, pp. 209–231, 2021.

[9] J. Kleinberg, S. Mullainathan, and M. Raghavan, "Inherent trade-offs in the fair determination of risk scores," arXiv preprint arXiv:1609.05807, 2016. doi: 10.48550/arXiv.1609.05807

[10] N. Bostrom and E. Yudkowsky, "The ethics of artificial intelligence," in Artificial Intelligence Safety and Security, Boca Raton, FL, USA: Chapman and Hall/CRC, 2018, pp. 57–69.

[11] R. Shokri, M. Stronati, C. Song, and V. Shmatikov, "Membership inference attacks against machine learning models," in Proc. IEEE Symp. Security Privacy (SP), 2017, pp. 3–18. doi: 10.1109/SP.2017.41

[12] M. Fredrikson, S. Jha, and T. Ristenpart, "Model inversion attacks that exploit confidence information and basic countermeasures," in Proc. 22nd ACM SIGSAC Conf. Comput. Commun. Security (CCS), 2015, pp. 1322–1333. doi: 10.1145/2810103.2813677

[13] C. Dwork, F. McSherry, K. Nissim, and A. Smith, "Calibrating noise to sensitivity in private data analysis," J. Priv. Confidentiality, vol. 7, no. 3, pp. 17–51, 2016. doi: 10.1007/11681878_14

[14] B. D. Mittelstadt, P. Allo, M. Taddeo, S. Wachter, and L. Floridi, "The ethics of algorithms: Mapping the debate," Big Data & Soc., vol. 3, no. 2, 2016, Art. no. 2053951716679679. doi: 10.1177/2053951716679679

[15] K. Reinhardt, "Trust and trustworthiness in AI ethics," AI and Ethics, vol. 3, no. 3, pp. 735–744, 2023. doi: 10.1007/s43681-022-00200-5

[16] F. Santoni de Sio and J. Van den Hoven, "Meaningful human control over autonomous systems: A philosophical account," Front. Robot. AI, vol. 5, p. 323836, 2018. doi: 10.3389/frobt.2018.00015

[17] G. Falco et al., "Governing AI safety through independent audits," Nat. Mach. Intell., vol. 3, no. 7, pp. 566–571, 2021. doi: 10.1038/s42256-021-00370-7

[18] C. E. Prunkl et al., "Institutionalizing ethics in AI through broader impact requirements," Nat. Mach. Intell., vol. 3, no. 2, pp. 104–110, 2021. doi: 10.1038/s42256-021-00298-y

[19] D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané, "Concrete problems in AI safety," arXiv preprint arXiv:1606.06565, 2016. doi: 10.48550/arXiv.1606.06565

[20] Y. Bengio et al., "Managing extreme AI risks amid rapid progress," Science, vol. 384, no. 6698, pp. 842–845, 2024. doi: 10.1126/science.adn0117

[21] E. Strubell, A. Ganesh, and A. McCallum, "Energy and policy considerations for deep learning in NLP," in Proc. 57th Annu. Meeting Assoc. Comput. Linguistics (ACL), 2019, pp. 3645–3650. doi: 10.18653/v1/P19-1355

[22] I. J. Goodfellow, J. Shlens, and C. Szegedy, "Explaining and harnessing adversarial examples," arXiv preprint arXiv:1412.6572, 2014. doi: 10.48550/arXiv.1412.6572

[23] X. Yi, Y. Bao, J. Zhang, Y. Qin, and F. Lin, "Integrating structural semantic knowledge for enhanced information extraction pre-training," in Proc. 2024 Conf. Empirical Methods Natural Lang. Process. (EMNLP), 2024, pp. 2156–2171. doi: 10.3390/app14166955

[24] T. Gu, B. Dolan-Gavitt, and S. Garg, "Badnets: Identifying vulnerabilities in the machine learning model supply chain," arXiv preprint arXiv:1708.06733, 2017. doi: 10.48550/arXiv.1708.06733

[25] Y. Bengio et al., "International AI safety report 2026," arXiv preprint arXiv:2602.21012, 2026. doi: 10.48550/arXiv.2602.21012

[26] W. D'Alessandro and C. D. Kirk-Giannini, "Artificial intelligence: Approaches to safety," Phil. Compass, vol. 20, no. 5, p. e70039, 2025.

[27] Y. Bai et al., "Constitutional AI: Harmlessness from AI feedback," arXiv preprint arXiv:2212.08073, 2022.

[28] S. Ghosh et al., "A safety and security framework for real-world agentic systems," arXiv preprint arXiv:2511.21990, 2025. doi: 10.48550/arXiv.2511.21990

[29] S. Russell, Human Compatible: AI and the Problem of Control. London, U.K.: Penguin, 2019.

[30] N. S. Bostrom, Superintelligence: Paths, Dangers, Strategies. Oxford, U.K.: Oxford University Press, 2014.

[31] I. Gabriel, "Artificial intelligence, values, and alignment," Mind. Mach., vol. 30, no. 3, pp. 411–437, 2020. doi: 10.1007/s11023-020-09539-2

[32] G. Irving, P. Christiano, and D. Amodei, "AI safety via debate," arXiv preprint arXiv:1805.00899, 2018. doi: 10.48550/arXiv.1805.00899

[33] N. Soares, B. Fallenstein, S. Armstrong, and E. Yudkowsky, "Corrigibility," in AAAI Workshop: AI and Ethics, 2015.

[34] T. Ord, The Precipice: Existential Risk and the Future of Humanity. New York, NY, USA: Hachette, 2020.

[35] A. Narayanan, "The limits of the quantitative approach to AI safety," Princeton University Center for Information Technology Policy, Princeton, NJ, USA, Tech. Rep., 2023.

[36] W. Wallach and C. Allen, Moral Machines: Teaching Robots Right from Wrong. Oxford, U.K.: Oxford University Press, 2008.

[37] M. Ienca, O. Buchholz, and E. Vayena, "AI ethical principles: The debate," in A Companion to Digital Ethics, Hoboken, NJ, USA: Wiley, 2025, pp. 101–112. doi: 10.1002/9781394240821.ch9

[38] A. F. Winfield and M. Jirotka, "Ethical governance is essential to building trust in robotics and artificial intelligence systems," Phil. Trans. R. Soc. A, vol. 376, no. 2133, 2018, Art. no. 20180085. doi: 10.1098/rsta.2018.0085

[39] G. E. Marchant and B. Allenby, "Soft law: New tools for governing emerging technologies," Bulletin of the Atomic Scientists, vol. 73, no. 2, pp. 108–114, 2017. doi: 10.1080/00963402.2017.1288447

[40] B. Wagner, "Ethics as an escape from regulation: From 'ethics-washing' to ethics-shopping?," in Being Profiled: Cogitas Ergo Sum, Amsterdam, The Netherlands: Amsterdam Univ. Press, 2018, pp. 84–89.

[41] European Commission, "EU Artificial Intelligence Act," Official Journal of the European Union, 2024.

[42] NIST, "Artificial intelligence risk management framework (AI RMF 1.0)," National Institute of Standards and Technology, Gaithersburg, MD, USA, Rep. NIST AI 100-1, 2023.

[43] M. Ienca, “Don’t pause giant AI for the wrong reasons,” Nat. Mach. Intell., vol. 5, no. 5, pp. 470–471, 2023.

[44] H. Roberts et al., "The Chinese approach to artificial intelligence: An analysis of policy, ethics, and regulation," in Ethics, Governance, and Policies in Artificial Intelligence, L. Floridi, Ed. Cham, Switzerland: Springer, 2021, pp. 47–79. doi: 10.1007/978-3-030-81907-1_5

[45] UK Government, "A pro-innovation approach to AI regulation," Dept. Sci., Innov. Technol., London, U.K., Policy Paper, 2023.

[46] T. AI, "OECD Digital Economy Papers," OECD Digital Economy Papers, 2023.

[47] A. A. Khan et al., "Ethics of AI: A systematic literature review of principles and challenges," in Proc. 26th Int. Conf. Eval. Assess. Softw. Eng. (EASE), 2022, pp. 383–392.

[48] P. Rastogi, "Role of AI in global partnership," J. Soc. Rev. Dev., vol. 3, no. Special 1, pp. 150–152, 2024.

[49] A. M. Barrett et al., "AI risk-management standards profile for general-purpose AI (GPAI) and foundation models," arXiv preprint arXiv:2506.23949, 2025.

[50] MEC Initiative, "The Minimum Ethics Code (MEC) v4.3: A universal technical standard for trustworthy AI," Zenodo, 2025. doi: 10.5281/zenodo.16938364

[51] S. Armstrong, A. Sandberg, and N. Bostrom, "Thinking inside the box: Controlling and using an oracle AI," Minds Mach.., vol. 22, no. 4, pp. 299–324, 2012. doi: 10.1007/s11023-012-9282-2

[52] X. Cheng et al., "Intelligent algorithm safety: Concepts, scientific problems and prospects," Bulletin of Chinese Academy of Sciences (Chinese Version), vol. 40, no. 3, pp. 419–428, 2024. doi: 10.16418/j.issn.1000-3045.20240720004

[53] R. Bommasani et al., "On the opportunities and risks of foundation models," arXiv preprint arXiv:2108.07258, 2021. doi: 10.48550/arXiv.2108.07258

Downloads

Published

2026-03-20

Issue

Section

Articles

How to Cite

Understanding Artificial Intelligence Ethics and Safety: Foundations, Challenges, and Governance. (2026). International Journal of Artificial Intelligence and Science, 3(1), 63-83. https://doi.org/10.63158/IJAIS.v3i1.61