Hidden Costs of LiteLLM's Health Checks
How LiteLLM model health checks generated unexpected external-model spend, and how we kept checks for local models only.
LLM inference, AI infrastructure, and distributed systems
How LiteLLM model health checks generated unexpected external-model spend, and how we kept checks for local models only.
An update on OCI model volumes, inference-aware routing, KServe, llm-d, multi-node inference, GPU scheduling, and the remaining gaps.
Run Kubernetes with NVIDIA GPU support inside WSL2 on a laptop, from k3s setup and validation to the limits of local GPU workloads.