Starts a vLLM inference server exposing the model behind an OpenAI-compatible API on http://localhost:8000, with endpoints like /v1/chat/completions and /v1/models. The model ID is a Hugging Face repo name and is downloaded on first start, which can take a while. Swap in any other repo name to serve a different model.
Looking for more? Search all 7,657 commands — works offline, in English or Spanish, and fixes typos.