Captions widen your audience considerably — people watching without sound, people in noisy places, people who need them. They are usually priced per minute of audio, which for a weekly event adds up quickly.
Why they normally cost money
Most captioning sends your audio to a cloud service for transcription and charges for the processing. Accurate, and you are paying for somebody else computer as well as sending your audio to a third party.
The local alternative
Speech recognition now runs well on ordinary computers. Processing the audio on the machine already running the broadcast means no per-minute charge, no API key, and no audio leaving the building.
It costs some processing. On a machine already encoding video that is a real consideration — use hardware encoding for the video so the processor has room for the transcription.
Burned in or as a separate file
Burned-in captions are drawn onto the picture. Everyone sees them and nobody can turn them off.
Sidecar files — VTT or SRT — sit alongside the recording and the viewer chooses. Better for the recording, and what platforms prefer for accessibility.
A common approach is burned-in live and sidecar files with the recording.
Getting the accuracy up
- Clean audio first. Captioning quality tracks audio quality almost exactly — a gated, close-miked voice transcribes far better than a distant one.
- Close microphones matter more here than anywhere else.
- Expect names, places and jargon to be wrong. Correct those in the recording afterwards.
- Do not promise perfect captions. Promise captions.
In Skynat Live
Captions are generated on the machine with no cloud key and no per-minute charge, and can be burned into the picture or written as VTT and SRT alongside the recording.
Because the audio console is in the same application, the gate and voice isolation that clean the sound also improve the caption accuracy — the transcriber hears the processed voice, not the raw room.