Automated Visual Prompting
摘要
Visual prompting (VP) is a parameter-efficient fine-tuning approach to adapting pre-trained vision models to solve various downstream image-classification tasks. This chapter presents AutoVP, an end-to-end expandable framework for automating VP design choices, along with 12 downstream image-classification tasks that can serve as a holistic VP-performance benchmark. The design space covers (1) the joint optimization of the prompts; (2) the selection of pre-trained models, including image classifiers and text-image encoders; and (3) model output mapping strategies, including nonparametric and trainable label mapping.