Agent Evals and Benchmarks
摘要
Agents feel impressive in demos; they earn trust with repeatable evals. This chapter gives you a Python-first harness to measure what matters: success on task suites, autonomy levels (how hands-off it really is), latency/cost trade-offs, and a process for incident post-mortems when runs go sideways.