System Prompt: all episodes

September 25, 2026

Before You Buy AI Hardware, Watch This

Audio version

In Episode 27 of System Prompt, Peter and Val break down inference runtime, the layer that sits between the model and the hardware serving it, and why runtime decisions matter before a business starts buying GPUs, DGX systems, or other local AI infrastructure.

The conversation covers Ollama, llama.cpp, vLLM, SGLang, quantization, context windows, KV cache, prefill, decoding, batching, speculative decoding, and parallelism. But the bigger question is business capacity. How many users will actually be active at once? Which departments need the most inference? When does demand spike? What should run locally, and what should stay in the cloud?

Using examples from a 150-person company to enterprise GPU deployments, Peter and Val argue that infrastructure should be sized from real usage data, not assumptions. A business may not need 150 concurrent sessions just because it has 150 employees,