Getting the text out of TikToks, Reels and Shorts

Burned in captions are pixels, not text, and often wrong. Pull the audio instead, work in batches, and check anything you plan to quote.

RIZZ AI LAB / 26 January 2026 / 5 min read

Short video is where most people now get told things, and it's the worst format to get text out of. The speech is fast, the captions burned into the picture are decoration rather than data, and the platforms don't want you leaving with a copy.

If you need the words, here's what actually works and what to expect.

Why the burned in captions don't help

Almost every TikTok, Reel and Short has captions on screen. They're part of the video image, drawn frame by frame. They aren't text. Selecting them is impossible because there's nothing to select.

They're also frequently wrong. Auto captions in the editing app get the words approximately right, then the creator styles them for rhythm rather than accuracy, splitting sentences across frames and dropping words that didn't fit. Reading them off the screen and typing them out gives you a version of what was said, not what was said.

The reliable route is the audio. Pull the audio, run speech to text, get the actual spoken words.

The three platforms behave differently

TikTok. Audio is straightforward to reach on public posts. Speech is usually fast and often over music, which is the main accuracy risk. Expect trouble with brand names and slang.

Instagram Reels. Public posts work. Private accounts don't, and nothing that requires a login will. If a Reel is on a private account, the only route is a screen recording you make yourself.

YouTube Shorts. These behave like ordinary YouTube videos, so caption tracks sometimes exist. When they do the text is cleaner than either of the other two.

Across all three the same rule holds: if the post is public, the audio is reachable. If it isn't public, no tool gets it without your own recording.

Sixty seconds of speech isn't much text

A one minute short is roughly one hundred and forty to one hundred and eighty spoken words. That's a third of a page. It's a small enough amount that people wonder whether extracting it is worth the trouble.

It's worth it in three situations.

Volume. Twenty shorts on the same topic is twenty minutes of watching and about three thousand words of text. Read as text it takes six minutes and you can see immediately which three of the twenty said anything the others didn't.

Citation. If you're going to repeat somebody's claim in your own work, you want the exact wording, not your memory of a video you watched at speed.

Search. A video you saved is a video you won't find again. A transcript you saved is searchable.

Doing it

With Drop, paste the post link. One credit per started minute, and a minute is the floor, so a forty second clip costs one credit and so does a fifty nine second one. Batch is where this gets cheap: fifteen shorts is fifteen credits.

What comes back is the transcript plus the core ideas and any lines worth keeping. For a short clip the ideas section is usually two or three sentences, which is honest, because a short clip usually contains two or three sentences of substance.

For a longer form version of the same job, getting the full transcript of any YouTube video covers the caption and speech to text tradeoff in more detail.

The screen recording case

Sometimes the words aren't spoken at all. A creator holds up a phone showing a workflow, or scrolls through a settings screen, or shows a prompt they used. The audio transcript will say look at this, which isn't useful.

For that, Drop reads the screen instead of the audio and returns what was visible: the tools named, the steps in order, the text that was legible on screen. That costs three credits rather than one because it's a different and heavier job. It's the right call for tutorial clips and the wrong call for talking head clips.

What to expect on accuracy

Short video is the hardest input in this whole category. Set expectations accordingly.

  • Music under speech reduces accuracy. Nothing fixes this fully.
  • Fast delivery is handled reasonably well by modern speech models.
  • Slang, product names and invented words will be guessed at.
  • Multiple voices in sixty seconds with no gaps will blur.

Check anything you intend to quote against the clip. That's thirty seconds of work and it's the difference between quoting somebody and misquoting them.

Where this ends up

Most people extracting short video are collecting for something: a piece of writing, a research question, a decision. The transcripts on their own aren't the output. Turning them into something publishable is covered in from transcript to post.

The short version: pull the audio, not the captions, do it in batches, and check the quotes before you use them.

Try it

Paste a link or drop a file in the box on the home page. One credit per started minute, and the first run on a new account is free.

Open the drop box

Read next