Multi-camera vision-based structural health monitoring of historic masonry minarets with LLM/VLM-assisted damage interpretation
摘要
This study proposes an LLM/VLM-orchestrated multi-agent framework for multi-camera vision-based Structural Health Monitoring (SHM) of historic masonry minarets. The main novelty of the framework is its region-aware and auditable decision-support strategy: camera-derived displacement anomalies, degradation in inter-sensor relationships, and VLM-based visual observations are preserved as traceable evidence streams and integrated at the decision layer rather than being merged through opaque feature-level fusion. The framework was experimentally validated on a scaled masonry minaret subjected to controlled shaking-table excitation. Multi-camera optical-flow tracking provided displacement time series for global and relation-based anomaly analysis, while region-specific inspection images were interpreted by a VLM as qualitative visual evidence. The global reconstruction-error pathway showed more frequent and persistent anomaly behavior during the damage-candidate phases compared with the reference condition. The relation-based pathway identified non-uniform degradation in inter-sensor consistency, with the strongest localization cue associated with the sensor region corresponding to the experimentally damaged area. Most importantly, the fused region-level risk map showed qualitative spatial agreement with the damage observed on the minaret model after the shaking table tests. These findings indicate that the proposed multi-agent framework can transform heterogeneous SHM evidence into interpretable regional risk priorities and reviewable reporting outputs for historic masonry structures, while maintaining a conservative expert-in-the-loop interpretation strategy.