Hidden Costs of LiteLLM's Health Checks
How LiteLLM model health checks generated unexpected external-model spend, and how we kept checks for local models only.
LLM inference, AI infrastructure, and distributed systems
I lead the architecture of Severstal’s shared platform for serving open-weight models on Kubernetes, vLLM, and NVIDIA GPUs.
Focus Control planes · Performance · Reliability · Observability · GPU capacity
How LiteLLM model health checks generated unexpected external-model spend, and how we kept checks for local models only.
Turn product scenarios, workload traces, latency SLOs, and failure requirements into an evidence-based GPU capacity plan.
An update on OCI model volumes, inference-aware routing, KServe, llm-d, multi-node inference, GPU scheduling, and the remaining gaps.
A production guide to monitoring vLLM 0.23.x with Prometheus and Grafana, including key metrics, PromQL, alerts, and runbooks.
Set up a modern C++ toolchain on Apple Silicon macOS with LLVM, CMake, Ninja, VS Code, Helix, and mise.
Run Kubernetes with NVIDIA GPU support inside WSL2 on a laptop, from k3s setup and validation to the limits of local GPU workloads.
Set up Helix for Python development with LSP, ty type checking, Ruff formatting and linting, plus a few editor refinements.
How vLLM improves LLM serving efficiency with PagedAttention, better KV-cache utilization, higher throughput, and steadier latency.
A concise uv cheat sheet for managing Python versions, environments, dependencies, tools, and scripts in one workflow.