Pathways, Limitations, and Future Directions for Embedding Values in Large Models
摘要
The rapid advancements in large-scale AI models have introduced significant concerns over value alignment, including biased outputs, privacy breaches, misinformation propagation, and diminished human autonomy. Building upon prior research on value embedding in AI, this paper offers a holistic assessment of value embedding strategies throughout the lifecycle of large models, from data collection, pre-training, post-training, and evaluation. We identify key technical barriers as well as regulatory hurdles that need to be addressed to ensure responsible AI development and deployment. To tackle these issues, we advocate a multi-pronged, socio-technical strategy that integrates more guaranteed AI safety approaches with more practical governance frameworks. By thoughtfully weaving value embedding into critical stages of AI development, we can steer large-scale models toward socially responsible and ethically aligned outcomes that reflect the diverse values of humanity.