whisper audio.mp3 --model base --word_timestamps True
Include word-level timestamps
Adds start and end time for every word to the JSON transcript, enabling highlight-style captions and forced alignment. Timestamps become far more granular than the default segment-level ones. The word data only appears in JSON or TSV output, not plain text.
Looking for more? Search all 7,657 commands — works offline, in English or Spanish, and fixes typos.