Hidden Costs of LiteLLM's Health Checks
How LiteLLM model health checks generated unexpected external-model spend, and how we kept checks for local models only.
Notes on software, systems, and things I learn along the way
How LiteLLM model health checks generated unexpected external-model spend, and how we kept checks for local models only.
Turn product scenarios, workload traces, latency SLOs, and failure requirements into an evidence-based GPU capacity plan.
A production guide to monitoring vLLM 0.23.x with Prometheus and Grafana, including key metrics, PromQL, alerts, and runbooks.