How to Plan LLM Inference Capacity for a Shared Platform
Turn product scenarios, workload traces, latency SLOs, and failure requirements into an evidence-based GPU capacity plan.
Notes on software, systems, and things I learn along the way
Turn product scenarios, workload traces, latency SLOs, and failure requirements into an evidence-based GPU capacity plan.