Limits the model's context window to 8192 tokens, trading longer inputs for lower memory use and faster prefill. Long prompts that exceed the limit are rejected instead of silently truncated. Reduce it further when VRAM is tight; raise it when memory allows.
Looking for more? Search all 7,657 commands — works offline, in English or Spanish, and fixes typos.