Limits how many sequences vLLM processes at once to 32, trading throughput for lower latency variance and predictable memory use. Each in-flight request holds KV cache space, so the cap matters when context windows are long. Raise it for chat workloads with many parallel users.
Looking for more? Search all 7,657 commands — works offline, in English or Spanish, and fixes typos.