Uncovering reliability blind spots in transportation AI: an edge-case-based validation approach for transportation systems
摘要
With artificial intelligence (AI) technology expanding into high-risk domains, criteria for evaluating its reliability and safety are required, as AI deployment may otherwise be constrained by future government regulation or societal backlash. In particular, verifying reliability in unexpected scenarios, such as edge cases, has emerged as a key challenge for implementing safe AI systems. This study introduces a supplementary validation approach to support the reliability testing of AI systems in high-risk domains. Considering the case of AI-based vehicle license plate recognition (VLPR), this study proposes edge-case-based validation approaches specifically designed for certification bodies. These approaches aim to address reliability validation needs that fall outside the well-defined operating regions covered by conventional testing. For VLPR systems, edge cases were identified as adverse-condition samples misrecognized by a baseline model and were then evaluated using a resolution-enhanced model. Under rainy conditions, the conventional recall metric reported 99.2%, whereas evaluating the same system on the extracted edge cases yielded 93.3%—a gap the aggregate metric did not make explicit. These findings highlight the potential for certification bodies to utilize this approach as a qualitative validation tool to uncover reliability blind spots in such systems. While the study focuses on VLPR systems, the proposed methodology is, in principle, transferable to other AI-based applications, although this remains to be tested. Future studies will aim to quantify edge case extraction techniques and examine their adaptability to diverse AI systems across various domains. This study provides preliminary, case-study insights that may contribute to improving the safety and reliability of AI systems in high-risk environments.