With the popularity of foundation models, recent years have witnessed a paradigm shift in deep learning from task-centric model design to task-agnostic representation learning and task-specific fine-tuning. Pretrained model representations are commonly evaluated extensively across various real-world tasks and used as a foundation for different downstream tasks. This chapter presents a solution called SynBench, as proposed in Ko et al. (What would gauss say about representations? probing pretrained image models using synthetic gaussian benchmarks. In: International Conference on Machine Learning (2024)), for assessing the quality of representations in a task-agnostic way. To circumvent the need for real-world data in evaluation, we explore the use of synthetic binary classification tasks with Gaussian mixtures to probe pretrained vision models and compare the robustness-accuracy performance on pretrained representations with an idealized reference. The approach offers a holistic evaluation, revealing intrinsic model capabilities and reducing the dependency on real-life data for model evaluation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Foundation Models Using Synthetic Datasets

  • Pin-Yu Chen,
  • Sijia Liu

摘要

With the popularity of foundation models, recent years have witnessed a paradigm shift in deep learning from task-centric model design to task-agnostic representation learning and task-specific fine-tuning. Pretrained model representations are commonly evaluated extensively across various real-world tasks and used as a foundation for different downstream tasks. This chapter presents a solution called SynBench, as proposed in Ko et al. (What would gauss say about representations? probing pretrained image models using synthetic gaussian benchmarks. In: International Conference on Machine Learning (2024)), for assessing the quality of representations in a task-agnostic way. To circumvent the need for real-world data in evaluation, we explore the use of synthetic binary classification tasks with Gaussian mixtures to probe pretrained vision models and compare the robustness-accuracy performance on pretrained representations with an idealized reference. The approach offers a holistic evaluation, revealing intrinsic model capabilities and reducing the dependency on real-life data for model evaluation.