LatentNeuroNet: A Text-Conditioned Stable Diffusion Framework for Reconstructing Visual Stimuli from fMRI
摘要
The human brain, among the most complex and mysterious aspects of the body, harbours vast potential for extensive exploration. Unravelling these enigmas, especially within neural perception and cognition, delves into the realm of neural decoding. Harnessing advancements in generative AI, particularly in the Image Processing domain, seeks to elucidate how the brain comprehends visual stimuli perceived by humans. The paper endeavours to reconstruct human-perceived visual stimuli using Functional Magnetic Resonance Imaging (fMRI). This fMRI data is then processed through pre-trained deep-learning models to recreate the stimuli. Introducing a new architecture named LatentNeuroNet, the aim is to achieve the utmost semantic fidelity in stimuli reconstruction. The approach employs a Latent Diffusion Model (LDM), emphasizing semantic accuracy and generating superior-quality outputs. Text conditioning within the LDM’s denoising process is handled by extracting text from the brain’s ventral visual cortex region. This extracted text undergoes processing through a Bootstrapping Language-Image Pre-training (BLIP) encoder before it is injected into the denoising process. In conclusion, a successful architecture is developed that reconstructs the visual stimuli perceived and finally, this research provides us with enough evidence to identify the most substantial regions of the brain responsible for perception.