System Prompt: all episodes

October 1, 2026

You've Bought AI hardware, now what?

Audio version

Buying AI hardware is only the beginning.

In Episode 28 of System Prompt, Peter and Val follow up on Before You Buy AI Hardware, Watch This by moving from planning into the actual hosting layer. Using GLM-5.3-Flash across two NVIDIA DGX Sparks with vLLM, they break down what the serving command is doing, which settings actually matter, and why hosting a model is very different from simply loading one.

The conversation covers tensor parallelism, distributed execution, RoCE and NCCL networking, Docker, model mounts, KV cache, context length, concurrency, Mixture-of-Experts execution, chat templates, reasoning and tool-call parsers, DFlash speculative decoding, prefill, decode, and live serving metrics.

Then they start the real cluster, bring the worker online before the head node, watch the model load across both Sparks, and run GLM-5.3-Flash through a coding workload. The live metrics