<p>Future prediction in visual systems is shifting from pixel-level frame generation toward latent-space forecasting, agent-centric reasoning, and planning-oriented decision support. This review summarizes this transition and analyzes how pretrained visual representations, temporal modules, graph-based interaction models, and language-conditioned planning reshape both modeling assumptions and evaluation protocols. The discussion first covers the move from direct video prediction to forecasting over compact semantic representations produced by frozen or adapted vision backbones. It then reviews recurrent, convolutional, transformer, state-space, diffusion, and world-model predictors for latent dynamics, followed by common evaluation pitfalls, including persistence shortcuts, static-scene dominance, metric mismatch, and insufficient horizon testing. The review further connects latent forecasting with multi-agent trajectory prediction, future-aware interaction reasoning, and instruction-conditioned inspection planning. The literature and case studies indicate that high feature similarity or low displacement error does not necessarily imply useful future understanding. Robust evaluation should therefore include simple baselines, horizon-wise metrics, dynamic-region or agent-level tests, multimodal uncertainty, and downstream decision performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From Pixels to Plans: A Review of Latent-Space Future Prediction

  • Ivan Malashin,
  • Vadim Tynchenko,
  • Dmitry Martysyuk

摘要

Future prediction in visual systems is shifting from pixel-level frame generation toward latent-space forecasting, agent-centric reasoning, and planning-oriented decision support. This review summarizes this transition and analyzes how pretrained visual representations, temporal modules, graph-based interaction models, and language-conditioned planning reshape both modeling assumptions and evaluation protocols. The discussion first covers the move from direct video prediction to forecasting over compact semantic representations produced by frozen or adapted vision backbones. It then reviews recurrent, convolutional, transformer, state-space, diffusion, and world-model predictors for latent dynamics, followed by common evaluation pitfalls, including persistence shortcuts, static-scene dominance, metric mismatch, and insufficient horizon testing. The review further connects latent forecasting with multi-agent trajectory prediction, future-aware interaction reasoning, and instruction-conditioned inspection planning. The literature and case studies indicate that high feature similarity or low displacement error does not necessarily imply useful future understanding. Robust evaluation should therefore include simple baselines, horizon-wise metrics, dynamic-region or agent-level tests, multimodal uncertainty, and downstream decision performance.