Zero-shot prediction of neighborhood health using multimodal large language models
摘要
Urban populations face growing multidimensional health risks, but scalable, data-efficient methods for neighborhood-level monitoring remain limited. Here, we introduce a zero-shot approach leveraging multimodal large language models (MLLMs) to predict neighborhood health outcomes without fine-tuning or labeled data. Using physical inactivity across four major U.S. cities as a case study, we demonstrate that task-specific prompt design significantly enhances ChatGPT’s performance, with combination of satellite imagery and socioeconomic indicators yielding optimal accuracy. City-wide implementations reveal that ChatGPT not only captures fine-grained spatial heterogeneity but also matches the predictive power of conventional supervised models, while circumventing the reliance on customized training sets and maintaining robustness across diverse urban contexts. By pairing publicly available data with general-purpose MLLMs, our framework provides policymakers and urban planners with an efficient, transferable tool for rapid health disparity assessment and intervention targeting. This work also underscores a paradigm shift in urban analytics—from purely data-driven modeling to knowledge-informed reasoning—and expands the frontier of MLLM applications in public health.