SDI embedded audio explained: groups, channels, embedders and de-embedders

0

In television production, audio does not always travel in its own cable. Most of the time it rides inside the video signal, embedded in the SDI. For an engineer coming from audio, that logic is disorienting: how do sixteen audio channels fit into a cable meant to carry pictures? Understanding embedded audio means understanding a large part of the plumbing of an OB truck, a studio floor or a master control room. This guide covers it all, from the principle to the field pitfalls.

What SDI embedded audio is

SDI (Serial Digital Interface) is the standard digital video link of professional television, over coaxial cable and BNC connectors. A video stream does not use all its bandwidth all the time: between active lines and frames, there are “ancillary data” spaces meant to carry side data. That is where audio sits. In other words, a single SDI cable carries the picture and up to sixteen digital audio channels, perfectly synchronous with it.

That coupling is the great strength of SDI embedded audio: picture and sound stay bound together from end to end. Where a separate audio install multiplies cables and risks of drift, SDI carries everything over one link, which simplifies the wiring of a truck or a studio.

Groups and channels: the structure

The sixteen audio channels are not loose: they are organized into four groups of four channels. Each group maps to two AES pairs, giving the following grid, universal in the trade.

GroupChannelsEquivalent
Group 1Channels 1 to 42 AES pairs
Group 2Channels 5 to 82 AES pairs
Group 3Channels 9 to 122 AES pairs
Group 4Channels 13 to 162 AES pairs

Embedded audio is generally at 48 kHz, the reference sample rate in video, and above all it is synchronous with the video signal: its clock is the picture’s clock. This group grid matters in operation, because you decide what to insert, extract and pass through in terms of groups.

Embedder and de-embedder: the two basic moves

Two operations structure the whole chain. The embedder (or audio multiplexer) inserts audio channels — from AES or analog inputs, or even a console — into the ancillary space of an SDI stream. The de-embedder (demultiplexer) does the opposite: it extracts one or more groups from the SDI to output them as AES, analog or to an audio network.

In real life, these functions take several forms: small dedicated converters at the edge of the floor, cards in a video router frame, or functions built into the gear itself. Production switchers, cameras and recorders embed and extract audio constantly, often without anyone thinking about it. The key is to know, at each point in the chain, which group holds the sound you care about.

Field pitfalls

The first pitfall is synchronization. To embed audio into a video stream, that audio must be locked to the video clock; a non-synchronous source causes clicks or glitches at insertion. Hence the importance of clocking in any install, a topic that connects directly to PTP timing in audio-over-IP networks.

The second pitfall is lip-sync. Every video process — keying, conversion, scaling — adds delay to the picture. If audio is not delayed by the same amount, sound leads the image. Managing lip-sync means realigning audio to the video latency, at embedding as at transmission. Finally, watch levels and broadcast loudness: a misconfigured de-embedder, and the whole level calibration goes sideways before the signal even reaches air.

From SDI to IP: when the streams split

The move to audio-over-IP changes things. With SMPTE ST 2110, picture and sound no longer travel together: the standard separates the streams, video on one side (ST 2110-20), audio on the other (ST 2110-30, AES67-compatible). You no longer de-embed a group: you subscribe to an audio stream on the network. It is a paradigm shift we saw arriving with gear like the Lawo Edge One stagebox, part of the broader migration described in our guide to audio over IP.

Still, SDI embedded audio is not going away: it remains everywhere on studio floors, in trucks and anywhere the infrastructure has not yet gone all-IP. Large broadcast consoles live astride both worlds and can de-embed SDI as readily as subscribe to a ST 2110 stream — true of surfaces like the Calrec Argo, the Lawo mc²96 or the SSL System T. Understanding embedded audio therefore stays essential, including to approach the IP transition calmly.

My take

SDI embedded audio is one of those topics you think belongs to video people, and that catches up with the sound engineer the moment they step into TV. The idea of sending sixteen channels through a picture cable has a certain elegance: one link, picture and sound bound together, wiring that breathes. The price is two disciplines you can never drop, clocking and lip-sync, which punish any sloppiness. IP and ST 2110 split the streams again for flexibility, but for a few more years, reading groups and driving a de-embedder stays a trade reflex. The real question is not “SDI or IP?” but how far your infrastructure still mixes the two.

Share.

About Author

After 20+ years in professional audio: live sound engineering, studio technical direction (Deep Forest, Pierre Jacquot), head of digital marketing at Playback.fr. A first-hand witness to the analog-to-digital shift, I track the whole audio landscape and break it down here — no fluff.