Thain at Scale
摘要
Chapter 10 hardens Thain through a second release cycle across three sprints. Sprint 2.1 establishes request-level visibility into token usage, latency, and estimated cost. Sprint 2.2 introduces governed optimization through response compaction and bounded caching, accepted only after passing an LLM-as-judge regression gate. Sprint 2.3 adds centralized reliability enforcement with timeout budgets, bounded retries, and controlled chaos scenarios for OpenAI, Azure AI Search, and Cosmos DB. The chapter concludes with a three-plane reference architecture that maps every book pattern to a production deployment, and seven enduring principles for agentic system design.