Guardrails for LLMs: A Comprehensive Review and Case Study on Bias Mitigation for Responsible AI
摘要
Large Language Models (LLMs) often encode latent biases that are difficult to eliminate through training or design alone. Like humans requiring education to regulate instinctive behaviour, LLMs need external safeguards to guide their outputs responsibly. Guardrail techniques have thus become essential for promoting fairness, safety and cultural sensitivity while preserving the model’s original capabilities. However, many existing guardrails reflect Western-centric norms, which can overlook linguistic diversity and national contexts, especially in smaller or multicultural societies. A case study on Māori language shows how rigid safeguards can distort meaning and suppress Indigenous perspectives. Responsible AI requires not only technical control but also culturally grounded approaches that are informed by local languages, values and policy environments.