Remaking Content Moderation
摘要
If generative hate slips through the cracks of moderation, what kinds of approaches might be more effective? Working closely with a software developer over several months, we conduct an array of experiments with AI vision models, from clustering ‘similar’ imagery to assembling machinic descriptions of hateful memes based on automatically generated captions. The aim here is not to discover a silver bullet solution, but to explore approaches that combine the strengths of computer vision with the strengths of human insight. We suggest that the descriptive abilities of AI models might be augmented with the contextual abilities of experts to combat the rising tide of hate on platforms.