Lists the models Ollama has loaded into memory right now, with their size, how long they have been resident, and how much VRAM or RAM they occupy. A model stays loaded for a few minutes after its last request unless `ollama stop` is used. This is the first place to look when a model feels slow or memory is tight.
Looking for more? Search all 7,657 commands — works offline, in English or Spanish, and fixes typos.