How to get the full transcript of any YouTube video
Captions, automatic captions and real speech to text are three different things. Which one you get explains most of the frustration, and only one of them works when captions are off.
RIZZ AI LAB / 12 January 2026 / 6 min read
Most people looking for a YouTube transcript want one of three things: to read instead of watch, to search inside a long video for one sentence, or to keep the material as text they can quote later. All three need the same starting point, which is the full spoken text with nothing missing.
The three ways text comes out of a video
There are only three real methods, and knowing which one you're getting explains most of the frustration people have.
Creator uploaded captions. Some channels upload a proper caption file. These are the best case: correct spelling, correct names, sentence punctuation. They're also rare outside large channels and anything produced by a broadcaster.
Automatic captions. YouTube generates these with speech recognition. They exist on most videos in major languages. They're usually accurate on ordinary speech and unreliable on names, technical terms and anything said over music. They also arrive as a stream of timed fragments with no sentence breaks, which is why pasting them into a document gives you a wall of lowercase text.
Fresh speech to text. The audio is pulled and run through a speech model. This is what you want when captions are switched off, when the video is in a language YouTube didn't caption, or when the automatic captions are too rough to work with. It costs more compute, so free tools tend to avoid it.
Doing it by hand
If you only need one video and it has captions, the manual route works. Open the video, use the three dot menu under the player and choose the transcript option. A panel opens beside the video with timestamped lines. There's a setting inside that panel to hide timestamps, which is the step most people miss. Turn it off, select the whole panel, copy.
What you get is readable but not clean. Speaker changes are invisible. Sentences run together. The text carries the caption line breaks, so a paragraph arrives as fifteen short lines. If you're pasting into a document you'll spend a few minutes joining lines back together.
The manual route also fails on the cases people care most about: private videos you've access to, videos with captions disabled, and non English videos with no caption track.
When the transcript is only step one
Here's the honest part. Very few people want a raw transcript for its own sake. A ninety minute interview is roughly fourteen thousand words. Nobody reads fourteen thousand words to find out whether the guest said anything useful.
What people actually want is one of these:
- the argument, in six sentences
- the four or five claims that carry the argument
- the two or three lines worth quoting, with the exact wording preserved
- a searchable copy kept somewhere they'll look again
That means the transcript is a raw material, not the product. Any workflow that stops at the transcript leaves you with the same problem you had before, just in a different format.
What Drop does with a YouTube link
Paste the link. Drop pulls the spoken text, and if there's no usable caption track it runs real speech to text on the audio rather than giving up. Then it reads the transcript and returns the core ideas in plain sentences, the quotes with their original wording intact, and the full transcript underneath so you can search it or download it.
The cost is one credit per started minute of material. A twelve minute video is twelve credits. A ninety minute podcast is ninety. Credits don't expire and the first run on a new account is free, so you can check the output quality on something you already know well before deciding anything.
If the link is public, the summary also gets recorded in the public Knowledge Ledger, which means somebody else who drops the same video later can reuse the work instead of paying for it again. Private sources and uploaded files never go there.
Things that go wrong, and what they mean
The transcript comes back very short. Usually the video has almost no speech, or the speech is buried under music. Drop treats anything under forty words as a thin result, doesn't record it publicly, and doesn't charge for it.
Names are spelled wrong. Speech recognition guesses at proper nouns. If you're quoting somebody, check the spelling of names against the video description before publishing. No system gets this reliably right.
Two people talk over each other. Overlapping speech is the hardest case in the field. Expect the crossover section to be approximate.
The video is in a language you don't read. Ask for the ideas in your own language while keeping the quotes in the original. A quote translated is no longer a quote.
A short workflow that holds up
- 1Drop the link and read the core ideas first. Ninety seconds tells you whether the video is worth more of your time.
- 2If it's, read the quotes. These are the parts you'd have highlighted if you had watched with a pen.
- 3Only then open the full transcript, and only to search it for the specific thing you came for.
- 4Save it. A transcript you can't find again in three weeks wasn't worth extracting.
That last step is where most people lose the value. There's a longer piece on the habit side of this in watch less, keep more, and if your material is mostly long form audio rather than video, turning a podcast episode into notes covers the differences.
The short answer
If the video has good captions and you need it once, use the transcript panel in YouTube and clean it up by hand. If you need it often, need the videos that captions don't cover, or want the ideas rather than the text, run it through something that does the speech to text and the reading in one pass.