Lets vLLM consume up to 90 percent of GPU memory for the KV cache, the buffer that stores conversation context, instead of the conservative default. Raise it on dedicated inference boxes to fit longer contexts; lower it when the GPU also runs other workloads. Combined with --max-model-len it is the main VRAM tuning knob.
Looking for more? Search all 7,657 commands — works offline, in English or Spanish, and fixes typos.