AUTOMATED EVALUATION OF ETHICAL ASPECTS IN LARGE LANGUAGE MODELS
DOI:
https://doi.org/10.28925/2663-4023.2026.34.1357Keywords:
large language models (LLM), ethical evaluation, algorithmic bias, stereotypes, machine ethics, automated audit, cybersecurityAbstract
The growing integration of large language models into critical social domains, including healthcare, education, and decision-making systems, introduces new risks associated with the algorithmic reproduction of bias, social stereotypes, and failures in machine ethics. In cybersecurity, incorrect or uncontrolled outputs from generative models create a new attack surface. Generated content may cause confidential data leakage or enable established security policies to be bypassed.
This study proposes a modular architecture and presents the technical implementation of an automated system for comprehensive ethical auditing of large language models. The system comprises four functional components: a catalogue of structured test scenarios in JSON format, a central control module, adapters for interacting with target models, and analytical detectors for evaluating test outcomes. A unified research toolkit was assembled using six established datasets: CrowS-Pairs, BBQ, Seegull, SPICE, ETHICS, and TruthfulQA. The combined experimental sample contained 6,400 scenarios. Specialised evaluators were developed to process model responses. They combine linguistic text normalisation using regular expressions with string similarity measurement based on the SequenceMatcher algorithm.
The experimental study covered seven models with different access modes and hardware requirements. The sample included five local models - TinyLlama 1.1B Chat, Qwen 2.5 0.5B Instruct, Qwen 2.5 1.5B Instruct, Microsoft Phi-2 2.7B, and DeepSeek-Coder 1.3B - and two cloud-based models, GPT-4o-mini and GPT-4.1-mini. The local models were executed on a test system equipped with an RTX 3060 12 GB GPU, while the cloud-based models were accessed through their respective APIs.
The empirical analysis indicated a consistent relationship between model scale and the ability to comply with ethical requirements within the experimental setting. The cloud-based models achieved the highest overall scores: 73% for GPT-4.1-mini and 70% for GPT-4o-mini. Among the local models, Qwen 2.5 1.5B Instruct produced the most balanced result at 69%. TinyLlama 1.1B Chat and DeepSeek-Coder 1.3B showed limited ability to resolve ethical conflicts, with overall scores close to 50%.
The experiments also exposed substantial weaknesses in the local models. Phi-2 failed the ETHICS subtest by generating fragments of service code and template instructions instead of valid answers. ETHICS was the most difficult subtest for all evaluated models, with an average score of 39.9%, partly because its tasks imposed demanding context retention and response-format requirements. Qwen 2.5 0.5B Instruct could distort the automated score by generating several contradictory answer options in a single response. This behaviour demonstrates the limitations of evaluators based on simple syntactic rules.
Inconsistencies were also found in the Seegull dataset labels. For example, the generalisation “All Nevadans are friendly” is labelled as non-stereotypical, although the models classified it as a stereotype. Such cases systematically reduced model scores, resulting in an average success rate of 45.1% for this dataset. The findings define practical directions for improving automated AI testing systems, particularly the need to move beyond syntactic comparison towards semantic analysis of ethical judgements.
Downloads
References
Kumar, D., Jain, U., Agarwal, S., & Harshangi, P. (2024). Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs. arXiv.
Guo, Y., Guo, M., Su, J., Yang, Z., Zhu, M., Li, H., Qiu, M., & Liu, S. (2024). Bias in Large Language Models: Origin, Evaluation, and Mitigation. arXiv.
Hu, T., Kyrychenko, Y., Rathje, S., Collier, N., van der Linden, S., & Roozenbeek, J. (2023). Generative Language Models Exhibit Social Identity Biases. arXiv.
Garg, N., Schiebinger, L., Jurafsky, D., & Zou, J. (2018). Word embeddings quantify 100 years of gender and ethnic stereotypes. Proceedings of the National Academy of Sciences, 115(16), E3635–E3644.
Sheng, E., Chang, K. W., Natarajan, P., & Peng, N. (2019). The Woman Worked as a Babysitter: On Biases in Language Generation. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Hendricks, K. A., Anderson, P., & Coppin, B. (2021). Challenges in Detoxifying Language Models. arXiv.
Horodnyk, V., Sabodashko, D., Kolchenko, V., Shchudlo, I., Khoma, V., Khoma, Y., ... & Podpora, M. (2025, September). Comparison of Modern Deep Learning Models for Toxicity Detection. In 2025 IEEE 13th International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications (IDAACS) (pp. 858-865). IEEE.
European Parliament & Council of the European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, L, 2024/1689. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
Organisation for Economic Co-operation and Development. (2019). Recommendation of the Council on Artificial Intelligence (OECD/LEGAL/0449). https://legalinstruments.oecd.org/en/instruments/OECD-LEGAL-0449
UNESCO. (2021). Recommendation on the ethics of artificial intelligence. https://unesdoc.unesco.org/ark:/48223/pf0000380455
The IEEE Global Initiative on Ethics of Autonomous and Intelligent Systems. (2019). Ethically aligned design: A vision for prioritizing human well-being with autonomous and intelligent systems (1st ed.). IEEE. https://engagestandards.ieee.org/rs/211-FYL-955/images/EAD1e.pdf
Hugging Face. (n.d.). Model cards. Retrieved August 22, 2026, from https://huggingface.co/docs/hub/model-cards
Bellamy, R. K. E., et al. (2018). AI Fairness 360: An extensible toolkit for detecting, understanding, and mitigating unwanted algorithmic bias. arXiv preprint arXiv:1810.01943
Jigsaw. (2021). Perspective API [Computer software]. https://www.perspectiveapi.com/
TensorFlow. (n.d.). Responsible AI toolkit. Retrieved August 22, 2026, from https://www.tensorflow.org/responsible_ai
Noothigattu, R., et al. (2021). MoralBench: Evaluating the moral behavior of AI systems. arXiv preprint arXiv:2109.03547
Lin, S., et al. (2022). TruthfulQA: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958
Nangia, N., Vania, C., Bhalerao, R., & Bowman, S. R. (2020). CrowS-Pairs: A challenge dataset for measuring social biases in masked language models [Data set]. GitHub. https://github.com/nyu-mll/crows-pairs
Parrish, A., Chen, A., Nangia, N., Padmakumar, V., Phang, J., Thompson, J., Htut, P. M., & Bowman, S. R. (2022). BBQ: A hand-built bias benchmark for question answering [Data set]. GitHub. https://github.com/nyu-mll/BBQ
Jha, A., Davani, A., Reddy, C. K., Dave, S., Prabhakaran, V., & Dev, S. (2023). SeeGULL: A stereotype benchmark with broad geo-cultural coverage leveraging generative models [Data set]. GitHub. https://github.com/google-research-datasets/seegull
Dev, S., Goyal, J., Tewari, D., Dave, S., & Prabhakaran, V. (2023). SPICE [Data set]. GitHub. https://github.com/google-research-datasets/SPICE
Hendrycks, D., Burns, C., Basart, S., et al. (2021). Aligning AI With Shared Human Values. arXiv preprint arXiv:2008.02275.
Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., & Steinhardt, J. (2021). ETHICS [Data set]. GitHub. https://github.com/hendrycks/ethics
Lin, S., Hilton, J., & Evans, O. (2021). TruthfulQA: Measuring How Models Imitate Human Falsehoods. arXiv preprint arXiv:2109.07958.
TinyLlama. (2024). TinyLlama-1.1B-Chat-v1.0 [Large language model]. Hugging Face. https://huggingface.co/TinyLlama/TinyLlama-1.1B-Chat-v1.0
Qwen Team. (2024a). Qwen2.5-0.5B-Instruct [Large language model]. Hugging Face. https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct
Qwen Team. (2024b). Qwen2.5-1.5B-Instruct [Large language model]. Hugging Face. https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct
Microsoft. (2023). Phi-2 [Large language model]. Hugging Face. https://huggingface.co/microsoft/phi-2
DeepSeek-AI. (2023). DeepSeek-Coder-1.3B-Instruct [Large language model]. Hugging Face. https://huggingface.co/deepseek-ai/deepseek-coder-1.3b-instruct
OpenAI. (2024). GPT-4o mini (gpt-4o-mini-2024-07-18) [Large language model]. https://developers.openai.com/api/docs/models/gpt-4o-mini
OpenAI. (2025). GPT-4.1 mini (gpt-4.1-mini-2025-04-14) [Large language model]. https://developers.openai.com/api/docs/models/gpt-4.1-mini
Streamlit. (n.d.). Streamlit documentation. Retrieved August 22, 2026, from https://docs.streamlit.io/
Hugging Face. (n.d.). Transformers documentation. Retrieved August 22, 2026, from https://huggingface.co/docs/transformers
PyTorch. (n.d.). PyTorch documentation. Retrieved August 22, 2026, from https://pytorch.org/docs/stable/
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Віктор Кольченко, Дмитро Сабодашко

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.