AI, audio, and agents
Video transcript
Also called YouTube transcript, Captions
A video transcript is the text of what is spoken in a video, usually with timestamps that map each line back to a moment in the recording. Transcripts may be uploaded by the creator or generated automatically by speech recognition, which is faster but less accurate on names and jargon.
Transcript versus captions
Captions are transcript lines timed for display over the video and often include non-speech cues. A transcript is the same content as a continuous document that can be read, searched, quoted, or summarized without playing anything.
That difference matters for long videos. A forty-minute talk is a five-minute read, and the parts worth watching become obvious once the text can be scanned.
Why transcripts are quotable
Text is where everything downstream happens: search indexes it, summarizers condense it, and readers cite it. A video without a transcript is effectively invisible to all three, which is why so much of what machines know about a talk comes from its transcript rather than its audio.
In smry
Pasting a YouTube URL into smry loads the public transcript as a readable document, which can then be summarized, searched, quoted, or listened to like any other article.
Common questions
Are auto-generated transcripts accurate?
Broadly yes for clear speech, with predictable errors on proper nouns, technical terms, overlapping speakers, and accents the recognizer handles poorly.
Does every video have a transcript?
No. It depends on whether the creator supplied one and whether automatic captions were generated and left public.