Orchestrate distributed LLM inference using KServe, Ray Serve, and GPU bin-packing on Kubernetes for optimized latency and maximum hardware utilization.