A fast model isn't a fast API. Here's how we worked backward from a 600ms p99 budget — Cloudflare Workers at the edge, WireGuard straight to the GPU — without giving up Kubernetes, global capacity, or fault tolerance.