ML Models Deployed: What Is Next?
摘要
Machine learning (ML) models augment many software applications across consumer, industrial, and military domains. It is known that the performance of ML models degrade over time. Model degradation can lead to incorrect predictions with potentially risky or harmful consequences. The state-of-the-art approach for model degradation is the use of model monitoring applications that detect and report model anomalies, based on model metrics, such as feature importance or F1 scores. Such metrics and the reported anomalies are comprehensible for data scientists, but not so much for users without a data science background from the application domain (e.g., medical, manufacturing, finance). Since model metrics are disconnected from the application domain context, it can take up to several days for a data scientist to identify the root cause of an anomaly and how to resolve it. This chapter presents an interaction concept for model monitoring that was validated by end users. It will provide answer to the following questions: Q1: How to enable application domain experts (with limited or no AI background) to understand a reported model anomaly and its root cause so they can select and initiate a response action in a timely manner? Q2: How to enable application domain experts to validate that a selected and initiated responsive action was effective? Q3: What are the cross-validation results of the interaction concept with selected user involvement models (Endsley 1995; Endsley and Kaber 1999; Rasmussen 1974)?