Free SRT subtitle extractor for videos and recordings

Drop a video or a recording with English speech in it and get back an .srt subtitle file, written on your own device. No watermark, no account, no upgrade prompt. Nothing is uploaded, so there is no length cap to hit.

Drop your video or recording here

or click to browse. Any video or audio file with English speech in it. Stays on your computer.

You read it before you keep it

The whole transcript is on screen, with its timestamps, before anything is saved. If it misheard something, you see it there.

No signup, no watermark

No account, no trial, no upgrade prompt on the download. Ads pay for the site.

Nothing is uploaded

The speech recognition runs in your browser. After the one-time download of the recognizer, you can go offline and it still finishes.

English, done properly

One language on purpose: a model trained only on English beats a multilingual one of the same size at it. Other languages are not supported at all.

How it works

  1. Drop a video or audio file, or click the box to choose one.
  2. The first use downloads the speech recognizer — about 80 MB, once. Your browser keeps it for every visit after that.
  3. It finds the parts where someone is speaking and writes them down with timestamps, on your own device.
  4. Read the result, then download the .srt. It is named to match your video, so players pick it up by themselves.

What it is good at, and what it is not

It is built for clear English speech: screen recordings, lectures, podcasts, interviews, meeting recordings, videos of someone talking to a camera. On those the transcript is usually close enough that fixing the odd word takes less time than typing any of it.

It is weakest where every speech recognizer is weak: several people talking over each other, heavy accents, loud music under the voice, a phone across the room. It also only understands English — other languages come out as nonsense rather than being translated. The transcript is on screen before you download, so judge it by reading, not by hoping.

What you get

A plain .srt file: numbered subtitles, each with a start time, an end time and a line or two of text. Every video player and every subtitle upload form accepts it. Keep its name matching the video’s — this tool names it that way for you — and put the two files in one folder, and most players load it without being asked.

What it does not do

It does not translate, does not label who is speaking, and does not have an editor built in. It writes English speech down as English text with timestamps, and that is all. An .srt is plain text, so any text editor — Notepad, TextEdit, anything — can fix a misheard word in seconds.

Frequently asked questions

Is my video uploaded to a server?

No. The speech recognition runs on your own device. The first use downloads the recognizer itself — about 80 MB, saved in your browser for next time — and after that starts you can even go offline; it will still finish.

What languages does it understand?

English, and only English. That is a deliberate choice: the tool ships a model trained for English alone, which is smaller and better at English than a multilingual one of the same size. Other languages come out as nonsense, not as translations.

How accurate is it?

Good on clear speech — a screen recording, a lecture, a podcast, a meeting with decent microphones. It gets worse with heavy accents, crosstalk, background music and bad audio, like every speech recognizer. Read the preview before you download; that is why it is shown.

What file do I get back?

A standard .srt file, named to match your video. Keep the two files in one folder with the same name and most players — VLC, IINA, MPV, most TVs — load the subtitles automatically.

Does it translate, or tell speakers apart?

No. It writes down the English words it hears, with timestamps. No translation, no speaker names, no editing here — the .srt is plain text, so any text editor can fix a word afterwards.

Is there a length or size limit?

No fixed cap, because there is no upload. Long recordings simply take longer: the tool transcribes only the parts where someone is actually speaking, so an hour of meeting takes well under an hour, but it is not instant.

Why is the first run such a big download?

The recognizer is a neural network of about 80 MB, and doing the work on your device means downloading it once. Your browser keeps it, so the second file — even on another day — starts without downloading anything.

Does it add a watermark or need an account?

No watermark, no account, no trial, and no upgrade prompt on the download. The site is free and paid for by ads.

Other tools