Skip to content
GivenTool
النسخة العربية

Video to text (transcript)

Turn speech in a video into a timestamped transcript on your device, then export TXT, Word, SRT or VTT. Nothing is uploaded.

Loading tool…

Choose a video and this tool writes down what is said, line by line with timestamps. It uses OpenAI Whisper, an open speech-recognition model, running inside your browser tab: the audio track is decoded on your device and recognised by a model downloaded to your browser, so the video itself is never uploaded anywhere. The transcript appears while it is being made, you can correct any line, click a time to jump the player to that moment, search it, and export plain text, text with timestamps, a Word document or SRT/VTT subtitles.

Two model sizes are offered. In our test on a single 2.1 GHz processor core, Fast (Whisper tiny, about 44 MB to download once) got through a 5-minute English recording at about 4 times real speed, so a 10-minute video takes roughly 2.5 minutes; More accurate (Whisper base, about 81 MB) made noticeably fewer mistakes at about 2 times real speed, roughly 5 minutes for the same video. Speed depends on your processor, and the first run also waits for the download. Whisper handles 99 languages including English and Arabic, can translate speech from any of them into English, and works best on clear speech: music, crosstalk, heavy accents or a distant microphone all cost accuracy, so read the result through before publishing.

Long videos are processed in 30-second windows. When a sentence runs past the end of a window, the next window starts at the beginning of that sentence, so words are not cut in half or repeated. Files up to 60 minutes and 500 MB work on a desktop or laptop (Chrome or Edge recommended; keep the tab open until it finishes); on phones the limit is 20 minutes because mobile browsers have less memory.

How to use it

  1. Choose or drop a video file (MP4, MOV, WebM, MKV, AVI and more).
  2. Pick the spoken language, or leave it on automatic detection, and choose Fast or More accurate.
  3. Click Transcribe. The first time, the speech model downloads once; after that the lines appear as each part of the video is recognised.
  4. Click a timestamp to check a passage in the player and fix any line directly in the transcript.
  5. Export as plain text, text with timestamps, Word (.docx), or SRT/VTT subtitles, or copy the text.

Frequently asked questions

Is my video uploaded to a server?

No. The audio is decoded and transcribed inside your browser tab. The only download is the speech model itself (from this site, once), which your browser keeps for next time. You can disconnect from the internet after the model has loaded and it will still work.

How accurate is it?

On clear speech in a common language Whisper is usually good, but it is not perfect: names, numbers and technical terms are the usual errors, and background music, overlapping speakers or strong accents reduce accuracy further. The More accurate model helps. Always proofread before you publish or quote the transcript.

How long does it take?

In our test on a single 2.1 GHz processor core, the Fast model transcribed English speech at about 4 times real speed (a 10-minute video in roughly 2.5 minutes) and More accurate at about 2 times real speed (roughly 5 minutes). Slower computers and phones take longer, and the first run adds the model download. A progress bar shows the time remaining.

Which languages are supported?

The 99 languages Whisper was trained on, including English, Arabic, Spanish, French, German, Hindi, Urdu, Turkish and Chinese. Quality is best for widely spoken languages; choosing the language yourself is more reliable than automatic detection for short or noisy clips.

Can it translate?

Yes, into English only: tick "Translate the speech into English" and Whisper writes an English version of what is said in any supported language. For other target languages, transcribe first and translate the text separately.

Is there a length limit?

Up to 60 minutes and 500 MB on a desktop or laptop, and 20 minutes on phones and tablets, because the decoded audio and the model have to fit in the browser's memory. For longer recordings, cut them into parts with the video trimmer first.

Does it tell speakers apart?

No. Whisper writes what is said but does not label who said it. Add speaker names while editing the lines if you need them.