Current Location: > Detailed Browse

Personalized Alignment of Large Language Models and Its Impact on Moral Judgment

请选择邀稿期刊:
Abstract: With the advent of the human-machine symbiosis era, the ethical dilemmas and algorithmic biases of large language models (LLMs) have triggered widespread ethical concerns. Guiding artificial intelligence (AI) toward benevolence has thus become an urgent and challenging imperative. This research explores a personalized alignment approach based on the HEXACO personality model and examines its impact on the moral judgment of LLMs. Specifically, the study aims to verify whether LLMs can effectively achieve personalized alignment through prompting and to systematically evaluate how such alignment influences utilitarian tendencies in LLMs compared to humans across various moral dilemmas. By leveraging mature psychological frameworks, this research seeks to provide a scientific basis for constructing controllable and ethical AI alignment strategies.
Study 1 tested GPT-3.5, GPT-4, and ERNIE 3.5 using HEXACO-based personality prompts across six domains at high, low, and baseline levels, integrated with different gender roles. Manipulation checks were conducted using two distinct methods: a quantitative assessment using the HEXACO-PI-R scale and a qualitative personal story-writing task rated by independent human evaluators. Study 2 utilized a set of standardized moral dilemmas to assess utilitarian versus deontological choices in both LLMs and human participants. Human data were categorized into high and low personality groups for comparison, while the LLMs performed the same moral judgment tasks under various personality settings to identify shifts in decision-making patterns.
The results of Study 1 confirmed the feasibility of personalized alignment, demonstrating that LLMs can dynamically represent HEXACO personality traits through prompts. Among the LLMs tested, GPT-4 exhibited superior instruction-following capabilities and more distinct trait differentiation than the other LLMs. Findings from Study 2 revealed that personality alignment significantly alters the moral judgment of LLMs, though the impact varies across different models and personality domains. Specifically, traits such as Honesty–Humility, Agreeableness, and Conscientiousness were found to reduce utilitarian tendencies, leading to a preference for deontological responses. While some traits, particularly Honesty–Humility, showed stable and consistent effects between humans and AI, others displayed divergent or even opposite patterns, highlighting fundamental differences in their respective moral reasoning mechanisms.
The study reached three primary conclusions. First, LLMs are capable of exhibiting stable and distinguishable personality tendencies that can be activated through prompt-based alignment. Second, the influence of Honesty–Humility on moral judgment exhibits a consistent effect across humans and different LLMs, whereas other personality domains show inconsistencies. This suggests that while LLMs’ moral decision-making shares partial cognitive logic with humans, fundamental differences remain. Third, the personality metatrait of “Stability”—and particularly the Honesty–Humility domain—demonstrates a significant moral salience effect within the personalized alignment process. Based on these insights, this research proposes a personalized alignment framework utilizing the HEXACO model and personality metatrait theory to systematically shape the moral responses of AI, providing a psychological foundation for the development of safety, controllable and ethical AI systems. This framework emphasizes integrating psychological theories to mitigate ethical risks and ensure that AI behavior remains consistent with human values.

Version History

[V1] 2026-03-13 14:20:10 ChinaXiv:202603.00082V1 Download
Download
Preview
Peer Review Status
Awaiting Review
License Information
metrics index
  •  Hits3393
  •  Downloads1151
Comment
Share
Apply for expert review
  • Operating Unit: National Science Library,Chinese Academy of Sciences
  • Mail: eprint@mail.las.ac.cn
  • Address: 33 Beisihuan Xilu,Zhongguancun,Beijing P.R.China