Forces the model to run in float16 (half precision), halving memory use versus fp32 with negligible quality loss on modern GPUs. vLLM usually auto-detects a good dtype, so this flag mainly resolves mismatches with unusual weight formats. For quantized weights, prefer --quantization instead.
Looking for more? Search all 7,657 commands — works offline, in English or Spanish, and fixes typos.