Calls vLLM's chat endpoint with one user message, exactly like OpenAI's API, and prints the assistant's reply as JSON. The messages array carries the whole conversation history on multi-turn calls. Add "stream":true to receive tokens incrementally instead of one big response.
Looking for more? Search all 7,657 commands — works offline, in English or Spanish, and fixes typos.