Subtitles, Captions and Transcripts: The Difference Matters
Someone asked whether the video needs subtitles or captions, and it turned out those two words mean genuinely different things that determine what you actually need to build.
Someone on your team asked whether the sermon video needs "subtitles" or "captions," said with the assumption that these are two names for the same thing. They're not, and the difference actually changes which tool you reach for and what problem you're solving.
This gets confused constantly, including by video platforms themselves, which often use the words loosely or interchangeably in their own settings menus. Here's what each term technically means, why the distinction matters for church video specifically, and how a transcript fits into all of it.
Subtitles: assume the viewer can hear, but not understand the language
Subtitles were originally built for a specific situation: a viewer who can hear the audio perfectly well, but doesn't understand the language it's in. A foreign film with English subtitles is the classic example — the audience hears the actors' actual voices and reads a translation of what's being said.
Because subtitles assume the viewer can hear, they typically don't include non-speech sound information — a door slamming, music swelling, a phone ringing — since a hearing viewer picks that up from the audio itself. What subtitles add is purely the words, translated.
Captions: assume the viewer can't hear the audio at all
Captions were built for the opposite situation: a viewer who cannot hear the audio, whether because they're deaf or hard of hearing, or because they're simply watching without sound. Captions are almost always in the same language as the spoken audio (they're not translating, they're transcribing) and, done properly, include relevant non-speech information a hearing viewer would otherwise pick up — "[applause]," "[music playing]" — though sermon captions in practice usually skip this and focus purely on the spoken words.
| Subtitles | Captions | |
|---|---|---|
| Assumes viewer can hear? | Yes | No |
| Same language as audio? | Usually different (translated) | Usually the same |
| Includes non-speech sound cues? | Typically not | Ideally, yes |
| Typical church use | Rare — would require actual translation | Common — accessibility and muted viewing |
Where "burned-in" and "closed" fit in
Both subtitles and captions can be delivered two different ways: closed (a separate file the viewer can toggle on or off) or burned-in / open (rendered permanently into the video frame, always visible). This is a second, independent axis from the subtitle-versus-caption distinction — you can have closed captions, open captions, closed subtitles, or open subtitles. For church social media clips specifically, the practical recommendation is burned-in captions; the full reasoning is in sermon captions: why burned-in beats everything else.
If what you're actually after is getting captioned clips out the door without sorting through the terminology every week, Sermon Clips burns word-timed, same-language captions into every clip automatically.
Try it on your next sermon — 2 freeWhere transcripts fit
A transcript is different from both of the above — it's simply the text of what was said, with no timing information tying words to specific moments in the video. A transcript on its own can't be displayed as captions or subtitles until timing data is added, aligning each word or phrase to the moment it's spoken. Think of a transcript as the raw material, and captions or subtitles as one of several finished products it can become — you have the transcript, now what covers the others.
Why this terminology confusion happens so often
Part of the blame belongs to the platforms themselves. YouTube's own settings historically used "subtitles/CC" as a combined label, blurring a distinction that used to matter more clearly. Most everyday conversation treats the words as interchangeable, and for a lot of casual use, that's harmless — but once you're evaluating a captioning tool or explaining accessibility requirements to your team, the precise meaning matters, because "we need subtitles" and "we need captions" can point you toward genuinely different features.
A quick way to know which one you actually need
- 1If your audience can hear the sermon fine and speaks the language it's preached in, but might be watching muted or is deaf or hard of hearing — you need captions.
- 2If your audience genuinely doesn't understand the language the sermon is preached in and needs a translation — you need subtitles, which is a translation task, separate from captioning.
- 3If you just have a document of what was said with no timing information yet — you have a transcript, and it needs to be converted into one of the above before it functions as on-screen text.
For the overwhelming majority of church video — social clips, full sermon uploads, livestreams — the actual need is captions in the language the sermon was preached in, not subtitles in a different language. Knowing that up front saves a round of confused conversations with whoever's shopping for a tool.
A worked example: sorting out a real request
Say a staff member emails and says, 'can we get subtitles on the sermon for Mrs. Alvarez, she's hard of hearing.' Read literally, that's a request for captions, not subtitles — Mrs. Alvarez can presumably understand English fine, she just can't reliably hear the audio. Calling it subtitles doesn't cause a real problem here because everyone involved understands the intent, but it matters the moment you go looking for a tool or a vendor, because a service that specifically does subtitle translation into other languages is solving a different problem than the one you actually have, and may not offer the same-language, word-timed captioning you need at all.
Compare that to a second, genuinely different request: a church with a growing number of Spanish-speaking attendees asks whether the sermon can be shown with subtitles in Spanish while the pastor keeps preaching in English. That is an actual subtitles request — a translation task, not a transcription-and-timing task — and it needs a different kind of service or workflow entirely, one that produces translated text rather than same-language text.
Where accessibility law fits into this
Some churches ask about captioning specifically because of accessibility obligations, and it's worth being precise about what that actually requires rather than guessing. Legal captioning requirements in most places are built around the same-language definition — captions for people who can't hear the audio at all — not translated subtitles for people who don't understand the spoken language. A church weighing whether it needs to caption its content for accessibility reasons is, in the vocabulary of this article, asking about captions, and that's a separate legal question from whether it wants to offer translated subtitles as an additional, voluntary service to a multilingual congregation.
Auto-generated captions and where the terminology breaks down further
Automatic speech recognition tools — the kind built into YouTube, most editing software, and dedicated captioning services — generate same-language text and call it, depending on the tool, 'auto-captions,' 'auto-subtitles,' or just 'transcription.' Under the hood these are almost always doing the caption job (same-language text from audio), regardless of which word the interface uses. If a tool specifically offers translation into other languages as a distinct, separate feature from its same-language auto-generation, that's the signal it's actually doing subtitle work in the strict sense — otherwise, assume 'auto-subtitles' on most tools just means automatically generated captions.
How this plays out across formats: livestream, upload, and clips
| Format | What's actually needed | Term to use when asking a vendor |
|---|---|---|
| Full sermon livestream | Same-language captions for viewers who can't hear the live audio | Captions |
| Full sermon upload afterward | Same-language, closed or burned-in captions | Captions |
| Social media clips | Same-language, burned-in (open) captions | Burned-in captions |
| Multilingual congregation, spoken language unchanged | Translated on-screen text in a second language | Subtitles (translation) |
Frequently asked questions
- What is the actual difference between subtitles and captions?
- Subtitles assume the viewer can hear the audio but doesn't understand the spoken language, so they're a translation. Captions assume the viewer can't hear the audio at all — whether deaf, hard of hearing, or simply watching muted — so they're same-language text of what's being said, sometimes including non-speech sound cues.
- Does my church need subtitles or captions for sermon videos?
- Almost always captions, not subtitles. Unless you're specifically translating the sermon into another language for viewers who don't understand the language it was preached in, what you need is same-language text for viewers who can't hear the audio — that's the definition of captions.
- What's the difference between closed captions and burned-in captions?
- This is a separate distinction from subtitles versus captions. Closed captions live in a toggleable file the viewer can turn on or off; burned-in (open) captions are rendered permanently into the video frame and can't be turned off. Both can technically apply to either captions or subtitles.
- Is a transcript the same thing as captions?
- No. A transcript is just the text of what was said, with no timing information linking words to specific moments in the video. It has to be aligned to the audio's timing before it can function as captions or subtitles on a video.
- Why does YouTube call it "subtitles/CC" if they're different things?
- Many platforms, including YouTube, have historically used "subtitles" and "captions" as a combined or interchangeable label in their settings, even though the two terms technically describe different use cases. This is a common source of confusion, though it rarely causes a practical problem for casual viewing.
- Do sermon captions need to include sound effects like applause?
- Formally, captions are supposed to include relevant non-speech sound information for viewers who can't hear the audio at all, such as "[applause]" or "[music playing]." In practice, most church sermon captions focus purely on the spoken words and skip this, which is a reasonable simplification for most use cases.
- Can one video have both subtitles and captions?
- Yes — a video can technically offer same-language captions for viewers who can't hear it and translated subtitles in a different language for viewers who don't understand the spoken language, as separate toggleable options. Most churches only need the captions side, since translation into another language is a separate undertaking.
Church Video Tips, Weekly
Join 1,000+ church communicators getting actionable strategies for growing your congregation's digital presence — every week, free.
No spam, ever. Unsubscribe anytime.
Keep reading
Part of
Sermon Captions, Subtitles and TranscriptsHow churches caption sermon video and publish sermon transcripts — burned-in captions for social, accessibility for the hard of hearing, and transcripts that make an archive searchable.