ROLA: real-world object-centric learning with attention optimization
摘要
Object-centric learning aims to parse visual scenes into compositional and interpretable units, termed “slots”. Recent strides in this field, propelled by advanced architectures like Transformers, have significantly improved scene decomposition capabilities. However, despite these advancements, current methodologies struggle to adeptly manage the complexity and diversity inherent in real-world scenes, often yielding suboptimal decomposition and reconstruction results. A key limitation of existing frameworks is their reliance on DINO—either as a frozen encoder or as a source of encoded features serving as reconstruction targets for decoders—which hampers the efficacy of slot-to-image reconstruction and restricts adaptability to scenes with varying object counts. In this work, we present ROLA, a novel approach that harnesses objectness cues derived from established DINO-based object-centric learning or object discovery techniques to steer the optimization of Slot Attention. By introducing an innovative attention optimization mechanism, ROLA achieves state-of-the-art decomposition performance on real-world scenes, consistently surpassing DINO-based benchmarks. Featuring a flexible encoder-decoder architecture, ROLA incorporates a variational autoencoder (VAE)-based slot decoder that reconstructs input images from slots while fostering disentangled, object-centric representations. We substantiate the efficacy of our method through rigorous empirical assessments on benchmark datasets, including CUB-200-2011 (CUB-Birds), COCO, and PASCAL VOC. Our framework not only elevates the quality of scene decomposition and slot-to-image reconstruction but also enhances the interpretability of learned representations, advancing the state-of-the-art in object-centric learning.