Beyond the Lookup: Simulating Realistic User Uncertainty for the Evaluation of Conversational Agentic Recommenders
摘要
The evaluation of Conversational Recommender Systems necessitates robust protocols to measure utility and user satisfaction. While human-in-the-loop testing remains the gold standard, scalability and reproducibility constraints have driven the field toward User Simulators. However, current simulation paradigms predominantly utilize rigid templates or closed-source Large Language Models that exhibit idealized behaviors. These approaches fail to capture user ambiguity, resulting in benchmarks that overestimate system proficiency by assuming crystallized user intent. To address this limitation, we introduce a family of open-weight user simulation models capable of generalizing across diverse e-commerce domains. Leveraging Teacher-Student distillation, we operationalize three distinct behavioral stereotypes: Direct, Vague-Proactive, and Vague-Reactive. Our evaluation of state-of-the-art Agentic Generative Conversational Recommender Systems reveals a critical Robustness Gap: while agents perform proficiently with decisive users, performance collapses when facing passivity and ambiguity. These findings underscore the necessity of our scalable framework for rigorously stress-testing the next generation of conversational agents against realistic, non-cooperative user behaviors.