Cognitive Instability in AI Agents—Security Risks from Misjudgment, Hallucination, and Drift
摘要
Many discussions of AI security risks focus on bad actors—external threats, adversarial inputs, or insider misuse. Some of the most consequential risks stem from something far more mundane: the cognitive limitations of the AI itself. Hallucinations, overconfidence, misalignment, and context drift—these aren't rare edge cases. They are well-documented and, in many cases, expected behaviors of current-generation AI systems.