Why Human Captioners?

Short answer: Automatic (AI) captions are generated by speech-recognition software that guesses at words it has never heard. Human captioners are trained stenographers who write what was actually said — including the drug name, the acronym, the speaker with an accent, and the moment two people talk at once. For technical, medical, academic and legal content, human live captioning is still the only way to reach the accuracy that deaf and hard-of-hearing audiences, and the law, expect.

What automatic captions get wrong

Everyone has watched auto-captions turn a keynote into word salad. It isn’t random. Speech-recognition software fails in predictable places, and those places are exactly where the meaning lives:

Specialist vocabulary. Drug names, gene names, programming languages, product names, acronyms. The software has never seen “Kubernetes” or “pheochromocytoma” in this context, so it substitutes something it has seen. The sentence still looks like a sentence. It just isn’t the one the speaker said.

Proper names. Speakers, sponsors, institutions, cities. Auto-captions mangle them, and a misspelled sponsor on a 40-foot screen is a very public mistake.

Accents and non-native speakers. The international language of conferences is English, spoken by people from everywhere. Recognition accuracy drops sharply for accented speech — the audience that most needs captions gets the worst captions.

Crosstalk, Q&A and panel discussions. Two people talking, an audience question from the back of the room, laughter over a punchline. Software merges it or drops it. A captioner writes “[audience member]” and carries on.

Numbers, homophones and punctuation. “Two” or “too”? “Fifteen” or “fifty”? Where does the sentence end? Without punctuation and speaker changes, even accurate words become hard to follow.

Silence about its own mistakes. This is the dangerous one. Auto-captions never say “I didn’t catch that.” They produce confident, fluent, wrong text, and a reader who can’t hear the audio has no way to know.

What a human captioner actually does

In our work, we hear for others. A stenographer writes what they hear on a shorthand machine — stroking whole words and phrases at once, like chords on a piano — and translation software turns that steno into English on screen about three-quarters of a second (or less) later. At White Coat, our captioners write at speeds over 260 words per minute with accuracy very near 100%.

The speed is the visible part. The invisible part is preparation. Before your event we build a custom dictionary from your event website, program, speaker list, slides and glossary, so the specialist terms come out right the first time. During the event, the captioner is listening for meaning, not sound: they know the difference between “the patient was hypotensive” and “hypertensive” because they know medicine, not just phonetics. And when something goes wrong — bad audio, a speaker who wanders from the mic — they adapt, ask, and keep the audience in the room.

Why it matters more than it used to

Captioning was once a courtesy. Increasingly it’s a requirement. In the United States, the ADA and Section 504 require effective communication, and the Department of Justice’s 2024 rule on web accessibility holds state and local government bodies — including public universities and colleges — to WCAG 2.1 Level AA, which includes captions for live audio. In the UK, the Equality Act 2010 and the Public Sector Bodies Accessibility Regulations point the same way. Accessibility officers know that “we ran the auto-captions” is not a valid defense when the captions were wrong.

And beyond compliance: one in six people has some degree of hearing loss, and a large share of any international audience is reading English more easily than they hear it. Good captions are how those people get the content everyone else gets.

When automatic captions are fine

We’ll say it plainly: for a casual internal meeting, a personal video, or a rough draft transcript you intend to correct, auto-captions are a reasonable tool. Use them. The problem is using them where the stakes are high — a conference your sponsors paid for, a medical lecture a student is relying on, a public commencement, a legal or governmental proceeding — and where the people who depend on the captions can’t check them against the audio.

What our clients say about the difference

“Norma and her outstanding team came to our rescue after a suboptimal experience with a provider that was ruining everyone’s experience. Not only did they provide state-of-the-art live captioning for a very technical and challenging topic, but they were able to arrange their schedules and put everyone to work with just a few hours’ notice. The audience couldn’t believe the improvement.”
Josep Jaume
FullStackFest
“When White Coat Captioning started captioning my lectures in my second year, I noticed an amazing difference. All of the energy I had been spending on hearing the words, I could now spend learning the material... In fact, classmates who sat behind us often commented that they read over our shoulders to get information that they too had missed.”
Alycia Reppel, M.D.
UVM College of Medicine

Frequently asked questions

How accurate are AI captions compared with human captioners?

Accuracy for automatic captions varies enormously with audio quality, accent and vocabulary, and it’s lowest exactly where content is most technical. Trained stenographers deliver accuracy very near 100% on the same material, and the widely used benchmark for accessibility-grade captioning is 99% or better. And on the rare occasions when it is absolutely necessary because of speed, content, or a major distraction that may occur, our captioners are trained to summarize properly so that the audience gets the information it needs.

Isn’t AI getting better?

Yes — and for everyday speech in a quiet room it’s impressive. It still has no way to know your speaker’s name, your product’s name, or the term of art your field uses, and it still can’t tell you when it’s guessing. Those are the errors that matter.

What is CART, and is it the same as live captioning?

CART (Communication Access Realtime Translation) is the term the accessibility and higher-education world uses for live, word-for-word captioning by a stenographer. The industry mostly says “live captioning” now. Same service, same people. See our CART and Live Captioning Services page.

Can you fix auto-captions after the fact instead?

Not live, and honestly cleaning up bad auto-captions afterwards often costs more than getting it right the first time. We provide lightly edited transcripts after every event as part of our pricing.

Do you use any AI at all?

Our captioners use professional steno translation software, which is not speech recognition — it converts the captioner’s own strokes to text. Nothing in our workflow listens to your audio and guesses.

Want captions your audience will actually read?

Tell us about your event and we’ll send a quote within one business day.