Comparison of the Brain Visual Cortex and CNN Under Continuous Object Property Space
摘要
One important area of artificial intelligence (AI) is the alignment of deep learning models and the human brain. It’s been proved that Convolutional Neural networks (CNN) are candidate computational models of the primate brain visual stream. Recent advances in neuroscience showed that object properties, including animacy and size, played an important role in visual object representations. But the property representations in CNN and whether these representations are similar with those in the visual cortex are unclear. Using the THINGS dataset, we created a continuous object property space and applied voxel-wise encoding models to map properties to human brain fMRI data and CNN layer features. Dimension reduction of human object property ratings identified three key dimensions: grasp, animacy, and feeling, which organize an object property space. Then, model weight analysis produced property representation maps, highlighting the role of the higher-level visual cortex in property representations. Cluster analysis using these maps functionally parcellated the visual cortex into three regions, each associated with specific preferred objects and properties. Finally, property analysis in CNN indicated that property could predicted responses across all layers. In conclusion, representations of object continuous properties in CNN and the human brain are different, which enlightens direction of the future alignment study between deep learning models and the brain.