Generative Architectures
摘要
Generative AI images emerge from specific architectures and understanding these is key to grasping how models operationalize stereotypes. CLIP, the model that underlies Stable Diffusion, aligns images and text, treating them as interchangeable and setting up the logic of equivalence. This logic reproduces essentialist and reductive portrayals of people and cultures, such as White men as CEOs, Black men as criminals, and Jews as world-running financiers, collapsing entire cultures into caricature. This reduction aligns seamlessly with the logic of hate, which thrives on simplified representations and clear targets of blame. This chapter first steps through how generative image production works in relation to CLIP and Stable Diffusion and then unpacks the implications of using a contrastive framework versus a classification approach.