If attention scaled intelligence, can it scale morality?
摘要
Large language models (LLMs) have achieved remarkable gains in cognitive performance through attention mechanisms functionally inspired by human attention. This paper asks philosophically whether a comparable architectural insight could technically advance moral processing. We argue that current alignment techniques primarily shape outputs after representations have been formed. They therefore cannot realise, using Iris Murdoch’s loving attention approach, a just, reality-sensitive orientation toward others that operates at the level of representation. Drawing on Murdoch’s moral philosophy, we identify three substrate-neutral features of loving attention, namely locational, representational, and dispositional, that survive translation from human moral phenomenology to computational systems. On this basis, we propose that moral processing in LLMs should be studied not only as a problem of output control but also as a problem of representational architecture. We further formulate the loving attention geometry hypothesis, suggesting that if morally improved perception involves a systematic shift from ego distorted to more just representations, this transformation may leave detectable structure in LLM embedding and activation spaces. With this focus on Murdoch, the paper contributes a novel philosophical technical approach that links moral attention, representation learning, and AI alignment. It clarifies why architectural considerations matter for moral AI, distinguishes between the tractability and validation of moral geometry, and outlines design constraints for responsible research. We argue that exploring representational forms of moral attention is a promising and necessary direction for AI ethics research. If transformer attention has been able to scale intelligence, it is worth inquiring if representations and architectures can scale morality.