Improving Reliability in Multimodal Legal Interpretation of Traffic Signs
摘要
Recent advances in multimodal large language models (LLMs) have shown strong performance in visual recognition and semantic understanding, yet their reliability in safety-critical tasks such as traffic sign interpretation under realistic conditions remains unclear. In this paper, we systematically evaluate multimodal LLMs under unconstrained and constrained supervision settings and compare single-pass inference, multi-agent predictor-verifier refinement, and a confidence-gated hybrid architecture. Results show that gains under constrained settings primarily reflect restricted output spaces rather than deeper regulatory understanding. In the realistic unconstrained setting, performance remains unstable, especially in composite scenes requiring compositional reasoning over multiple interacting signs. Contrary to expectations, multi-agent self-refinement does not yield consistent improvements and can introduce instability through high-confidence but normatively incorrect feedback. In contrast, the proposed confidence-gated hybrid architecture that separates visual grounding from semantic reasoning substantially improves reliability.