SDI embedded audio explained: embedders, de-embedders and staying in sync

0

In television production, sound does not always travel on its own cables: more often, it rides inside the video signal. The SDI link (Serial Digital Interface), the coaxial BNC cable that ties cameras, switchers and control rooms together, carries the picture and, in the very same bytes, up to sixteen channels of audio. This is embedded audio, and knowing how to insert it, extract it and keep it in sync is a core broadcast-video skill.

What “embedded” means

An SDI stream reserves spaces between the active picture lines for ancillary data. Audio sits there as packets, organized into groups of four channels. An HD-SDI stream carries up to four groups, or sixteen channels, usually at 48 kHz locked to the video clock. The appeal is obvious: one cable, one connector, picture and sound travelling together from camera to control room with no risk of splitting them by mistake.

SDI-to-audio de-embedder that extracts audio from the SDI signal
Credit: Blackmagic Design

Capacity by SDI rate

InterfaceRateEmbedded audio
SD-SDI270 Mb/sup to 16 channels
HD-SDI1.5 Gb/sup to 16 channels (4 groups)
3G-SDI3 Gb/sup to 16 channels
12G-SDI12 Gb/s4K/UHD, 16 channels

Embedder and de-embedder: the two key moves

Two operations come up constantly. Embedding inserts audio channels (analog, AES/EBU or from a console) into an SDI stream that had none, or replaces the ones already there. De-embedding does the reverse: it extracts audio from the SDI to feed a console, a processor or a recorder. In the field these functions take the form of small converter boxes, or are built into production switchers, routing cards and multiplexers.

The typical floor scenario: audio embedded by the cameras arrives in the control room, a de-embedder sends it to the audio console, the engineer mixes, then an embedder reinserts the clean mix into the SDI stream headed for recording or air. This split-then-reinsert logic echoes, in another world, the way audio over IP separates the essences to route them independently.

Sync: the real trap

The number-one problem with embedded audio is lip-sync. Every video process — a corrector, a keyer, a standards converter — adds latency to the picture, often one or more frames. If the audio takes a shortcut, it arrives early: mouth and voice no longer match. The fix is to delay the audio to line it up with the picture, using the delay lines built into de-embedders and consoles. You always reason relative to the house sync reference (genlock): picture and sound must lock to the same clock, or the channels slowly drift and eventually click.

Another rule: embedded audio runs at 48 kHz tied to the video. Inserting a source at another rate without clean conversion produces artifacts. A serious embedder handles sample-rate conversion and respects the video-clock lock — the same timing demand that governs networks, as detailed in our piece on PTP and clocking in audio over IP.

Mapping the channels: the discipline that saves a live show

On a multicamera shoot, every source arrives with its own embedded channel layout, and nothing is more common than a lav mic on channel 3 of one camera and channel 1 of another. Without a clear convention, de-embedding turns into a treasure hunt. The field rule: define one channel map (say, program on channels 1-2, foldback and wireless on the rest), document it, and enforce it across all camera control rooms. De-embedders can remap channels on extraction; leaning on that saves you from reconfiguring an entire console because one camera broke the standard.

When channel counts explode — big productions, heavy multicam — SDI quickly hits its sixteen channels. You then move to MADI or AES/EBU for high-density audio tie-lines between the control room and peripherals, leaving SDI to carry picture plus local sound. Juggling these formats is really about seeing that embedded audio is just one link in a chain where each interface has its role.

SDI or IP: two worlds that meet

The shift to all-IP (SMPTE ST 2110) changes things: there, audio is no longer a prisoner of the picture, it flows as a separate stream on the network, which simplifies routing and mixing. But SDI is still everywhere — cameras, installed gear, whole production chains — and SDI/IP gateways are all around. Understanding embedded audio remains essential, if only to interface an SDI floor with an IP control room, or to route a signal to a contribution codec on a remote link.

Finally, embedded audio is not exempt from broadcast rules: the mix reinserted into the SDI must meet R128 loudness targets, and production intercom circuits often share the same infrastructure as the program channels.

My take

Embedded audio is one of those topics you assume is the video engineers’ problem, until the day an otherwise flawless mix reaches air half a second out of sync. The truth is that TV sound lives as much in the routing and the sync as in the console. The move to IP eases many of these constraints, but SDI is not going away tomorrow: as long as there are cameras and coax, you will need to extract, mix, realign and reinsert. Master those four moves and you spare yourself the call no sound engineer wants to get mid-broadcast.

Share.

About Author

After 20+ years in professional audio: live sound engineering, studio technical direction (Deep Forest, Pierre Jacquot), head of digital marketing at Playback.fr. A first-hand witness to the analog-to-digital shift, I track the whole audio landscape and break it down here — no fluff.