How to Plan LLM Inference Capacity for a Shared Platform
Turn product scenarios, workload traces, latency SLOs, and failure requirements into an evidence-based GPU capacity plan.
Production LLM inference, AI infrastructure, and distributed systems
Turn product scenarios, workload traces, latency SLOs, and failure requirements into an evidence-based GPU capacity plan.
An update on OCI model volumes, inference-aware routing, KServe, llm-d, multi-node inference, GPU scheduling, and the remaining gaps.
Set up a modern C++ toolchain on Apple Silicon macOS with LLVM, CMake, Ninja, VS Code, Helix, and mise.
Run Kubernetes with NVIDIA GPU support inside WSL2 on a laptop, from k3s setup and validation to the limits of local GPU workloads.
Set up Helix for Python development with LSP, ty type checking, Ruff formatting and linting, plus a few editor refinements.
A production guide to monitoring vLLM 0.23.x with Prometheus and Grafana, including key metrics, PromQL, alerts, and runbooks.
How vLLM improves LLM serving efficiency with PagedAttention, better KV-cache utilization, higher throughput, and steadier latency.
A concise uv cheat sheet for managing Python versions, environments, dependencies, tools, and scripts in one workflow.
Scan .NET dependencies for known NuGet vulnerabilities and automate checks in GitLab CI to improve supply chain security.
Convert FLAC to Apple Lossless (ALAC) with FFmpeg while preserving audio quality and metadata, then import the files into Apple Music.