Cost Management
摘要
Speed is nice; cheap and fast is better. This chapter is a practical kit for lowering inference and training spend without gutting quality: token budgets, hybrid routing (small-first), caching, distillation, and spot fleets + autoscaling. Everything is Python-first and meant to drop into a real service.