Mimicking the Mind’s Eye: AI-Driven Methodologies for Rorschach-Inspired Image Interpretation
摘要
The advancement of artificial intelligence has revolutionized the interpretation of images, following the principles of the Rorschach test. This paper presents a cutting-edge approach that merges image captioning models with a visual question answering (VQA) system to intricately extract and combine detailed image descriptions. Drawing inspiration from the Rorschach test, our study probes the complexities of human personality behavior. By employing Rorschach test cards as investigative tools, we delve into concealed personality traits, confront biases, and examine the challenges of multifaceted behavior. This research offers a fresh lens through which to view how people perceive and engage with their environment, thereby enriching the field of personality psychology and broadening our understanding in various settings. The study has significant implications for counseling, education, and interpersonal dynamics, aiding in the deep exploration of human behavior. Our methodology incorporates three distinct models—salesforce/blip, jaimin, and noamrot/FuseCap—for analyzing image descriptions. We perform a comparative evaluation of these models and user responses to generate personality-informed descriptions for each Rorschach test card, focusing on aspects like content, location, determinants, populars, and form quality. This technique reflects the human process of making sense of ambiguous images, venturing into new territories of AI-based interpretation, and augmenting the realm of psychological assessments.