How to summarize a two hour talk without watching it twice
Two hours of conference talk holds about twenty minutes of content. Convert first, triage in ninety seconds, then search the transcript for the one thing you came for.
RIZZ AI LAB / 2 February 2026 / 6 min read
A two hour conference talk contains roughly twenty minutes of content. This isn't a criticism of speakers. The format demands padding: the setup, the audience joke, the history section everyone already knows, the demo that goes wrong, the questions at the end from people who wanted to make a statement.
The job is finding the twenty minutes without sitting through the hundred and twenty.
Why watching at double speed isn't the answer
The standard advice is to play it at 2x. This halves the time and keeps the structure, which sounds reasonable and mostly isn't.
At double speed you lose the ability to notice. Comprehension holds up on familiar material and drops on anything new, which is exactly the material you sat down for. You also can't skim. Video has no skim mode. You can jump, but jumping is blind, so you land in the middle of a sentence and scrub around until you find a boundary.
Text has skim mode. That's the whole argument for converting first.
Step one, get the text
Whatever the source is, the first move is the same: turn two hours of speech into fourteen thousand words of text.
Conference talks are usually on YouTube or a conference site, and both are reachable by link. Internal recordings are usually a file, which means uploading it. Either way the output should be the full spoken text, not a summary someone else made.
Drop does this in one pass: the transcript plus the core ideas already pulled out of it. One credit per started minute, so a two hour talk is one hundred and twenty credits. Whether that's worth it depends entirely on whether the alternative is two hours of your time.
Step two, read the ideas before the transcript
The core ideas come back as a handful of plain sentences. Read them first. This is the triage step and it takes ninety seconds.
One of three things happens.
Nothing here. The talk restated things you know. Close it. You spent ninety seconds and two hours of viewing you were dreading is now resolved.
One thing here. There's a single claim or method worth understanding. Now you know what you're looking for, which turns the transcript from a wall of text into a search.
A lot here. Rare, but it happens. Now watching becomes worth it, and you can watch properly rather than at 2x, because you already know it'll pay.
Most talks land in the first two categories. That's the value: not summarising everything, but deciding fast which of the three you're in.
Step three, search the transcript for the specific thing
Once you know what you want, the transcript is a searchable document. Find the term, read the surrounding four hundred words, and you've the argument in full detail in about five minutes.
This is the part that a summary alone can't do. Summaries lose the reasoning. If the claim is that a technique cut deployment time by half, the summary tells you the claim and the transcript tells you the conditions under which it held. You usually need both, and you need them in that order.
Step four, keep the quotes, not the summary
If you're going to use the material, keep the exact lines. A summary you wrote isn't citable. A quote is. Drop separates these deliberately: the ideas are paraphrase, the quotes are verbatim, and it doesn't blur the two.
For anything you plan to publish, verify quotes against the recording. Speech recognition is good and not perfect, and proper nouns are where it fails.
The panel discussion problem
Panels are harder than talks. Four people, overlapping speech, no clear structure, and a moderator who asks a question every eight minutes that resets the topic.
Transcripts of panels are messy and speaker attribution is unreliable when voices overlap. What works better is treating the panel as a set of claims with the attribution held loosely, then checking who said the one you care about by jumping to that point in the recording. That's one lookup rather than two hours.
What this isn't good for
Some things don't survive conversion to text.
- Demos where the point is what happened on screen.
- Anything where slides carry the argument and the speaker only narrates.
- Performance, where how it was said is the content.
For the first two, reading the screen rather than the audio is the right tool. Drop does this as a separate mode at three credits per minute, and it returns what was actually visible rather than what was said about it.
The habit version
Doing this once saves an afternoon. Doing it as a default changes what you're willing to open, because the cost of checking whether a two hour talk is worth watching drops to almost nothing. That shift is the subject of watch less, keep more.
If the material is a long interview rather than a talk, the selection problem is slightly different and turning a podcast episode into notes covers it.
Summary
Convert first, triage in ninety seconds, search the transcript for the one thing, keep the quotes verbatim, and only watch when the text told you it's worth it. Most of the time it'll tell you it isn't, and that's the result you wanted.