Voice in a video edit: the processing chain most editors skip

0

The short version

  • A clean recording is not a finished voice: without a processing chain, dialogue stays dull, uneven and fragile as soon as music comes in.
  • The order of the stages matters as much as their settings — cleanup, high-pass and subtractive EQ, compression, de-essing, polish, then level calibration.
  • The loudness target depends on the destination: around -14 LUFS on YouTube, -16 LUFS for a podcast, -23 LUFS for European television.
  • Final Cut Pro, Premiere Pro and DaVinci Resolve can all do this work. What is missing is the method, not the tools.

The cut is locked, the grade is clean, the titles are in place — and the voice still sounds exactly like the file that came off the recorder. This has become so common that it now passes for normal: in most video production pipelines, dialogue travels all the way through without anyone really treating it. The recording was looked after, a decent microphone was placed at a sensible distance, levels were watched on set, and the audio job is assumed to be done.

It is not. A clean recording is raw material, not a result. Between the original file and a voice that holds its own against a music bed, in headphones as well as on a phone speaker, sits a processing chain of five or six stages. That chain is what is almost always missing, and this article walks through it stage by stage, tool by tool.

The problem is not the recording, it is the missing chain

A video editor works the image in successive layers: white balance, primaries, secondaries, masks, grain. Nobody would claim a shot is graded because the exposure was pushed half a stop. On audio, though, the gesture often stops there: pull the fader up until it can be heard, move to the next shot.

Spoken voice, however, is a signal with a very wide dynamic range. Between a syllable attacked at the start of a sentence and a breathed-out ending, 20 to 30 dB of difference is routine. No fader position fixes that: if the start of the sentence sits right, the end disappears under the music; if the end is audible, the start clips. A fader sets an average level, it does not shape a dynamic range.

The two symptoms that give away a missing chain are always the same. First, intelligibility collapses the moment a music bed arrives, because the voice was never cleared in the frequencies where music masks it. Second, perceived level jumps from one sequence to the next, because the takes come from different days, different rooms and sometimes different microphones, and no stage was ever tasked with bringing them to a common denominator.

The chain, in order

The voice processing chain, in order 1. Cleanup hiss, mains hum, room reverberation 2. HPF and EQ take away first, add afterwards 3. Compression two gentle stages beat one hard one 4. De-essing after compression, never before 5. Polish presence, air, sitting with music 6. Loudness LUFS target and true peak ceiling Each stage sets up the next one: cleaning after EQ means equalising noise, compressing before the high-pass makes the compressor trigger on frequencies nobody hears, and de-essing at the head of the chain is pointless since compression and the presence lift are what push sibilance forward in the first place.
The order of the stages shapes the result as much as the settings inside them.

That sequence is not an aesthetic convention, it is a logical constraint. A noise reducer analyses a spectrum: equalise before it and you hand it a spectrum that has already been bent. A compressor follows an envelope: leave in the sub-bass rumble of a lavalier rubbing against a shirt and it will trigger on events nobody can hear. A de-esser looks for sibilance: place it upstream and it works on a signal where sibilance is not yet a problem, and lets it straight back in at the next stage.

Stage 1 — clean before you treat

The iZotope RX 12 interface showing the spectral display of a dialogue track
iZotope

Four nuisances turn up on almost every dialogue take shot in real conditions: broadband hiss from a preamp pushed hard on an insensitive microphone, continuous background noise from ventilation or traffic, mains hum at 50 or 60 Hz with its harmonics, and room reverberation. The first three respond well to treatment. The fourth is the hardest, because it is correlated with the voice itself.

Work from the steady to the transient. A noise reducer trained on a profile — two seconds of silence taken from the same recording, not from another one — does most of the job with 6 to 9 dB of reduction, and no more. Past that, the ear hears the processing before it hears the voice: word endings turn metallic and the room ambience drops out abruptly between lines. A de-hummer set to the mains fundamental then clears 50 or 60 Hz and their multiples. Mouth clicks and heavy breaths are dealt with last, one by one, by hand or with a dedicated module.

Neural voice isolators are not meant to run at 100 %

Voice Isolation in Final Cut Pro and DaVinci Resolve Studio, Enhance Speech in Premiere Pro: these neural isolators are spectacular on a thirty-second demo and destructive across a forty-minute interview. Pushed all the way, they build a vacuum-packed voice with no ambience at all, dropped sentence endings and metallic artefacts on fricatives. The right approach is to dial them between 30 and 60 % and let the room live, topping it up with a neutral ambience bed under the dialogue if needed. We covered this whole family of tools in our round-up of real-time voice isolation tools.

Stage 2 — EQ starts with what you remove

The most widespread reflex, and the most counterproductive, is to open an equaliser and lift the top end until the voice cuts through. It does cut through, along with the hiss, the sibilance and the mouth noise. The right method is the opposite: take out what is in the way first, then lift far less than you expected to.

Where to work on a spoken voice High-pass 50 – 90 Hz Fundamental 90 – 200 Hz Mud 200 – 400 Hz Boxy, nasal 400 Hz – 1 kHz Harshness 1 – 2.5 kHz Presence 2.5 – 5 kHz Sibilance 5 – 9 kHz Air 9 – 16 kHz 80 100 200 400 800 1.6 k 3.2 k 6.4 k 12.8 k Blue zones are the ones you protect or lift very slightly. Grey zones are the ones where the useful move is subtractive. Logarithmic scale in hertz.
The map is a starting point, not a recipe: a low voice has a fundamental down around 85 Hz, a high one sits above 200 Hz.

Set the high-pass first, with a steep slope, between 70 and 90 Hz on a male voice and between 100 and 120 Hz on a female one. The reliable method is to raise the corner frequency until the voice starts to sound thin, then back off one notch. Anything below the fundamental adds nothing to the timbre and makes the following stages work for nothing.

Next comes the 200 to 400 Hz band, responsible for the cardboard-box effect. It is particularly loaded on lavaliers worn under clothing and on close-miked takes, where proximity effect swells the low mids. A wide bell with 2 to 4 dB of cut is almost always enough. Between 400 Hz and 1 kHz live nasality and room colouration: there you need a narrow bell, swept until you find the exact spot, and a surgical cut of 3 to 5 dB.

Only after that cleanup do you lift presence, between 2.5 and 5 kHz, with a wide bell and 2 to 3 dB at the very most. The air shelf above 10 kHz comes last, gently, and never on a noisy take — it would bring the hiss up with everything else. The behaviour of each filter type is broken down in our guide to equalisation.

Stage 3 — two gentle compressions beat one brutal one

The Fairlight page in DaVinci Resolve, with its audio timeline and mixing buses
Blackmagic Design

Compression is the stage video editors get wrong most often, because they treat it as a level corrector rather than an envelope shaper. A single compressor asked to absorb 10 to 12 dB of dynamic range produces an instantly recognisable result: the voice pumps, breaths jump to the front, and the room floor rises and falls with every sentence.

The fix is to split the work across two stages. The first one, slow, evens out sentences against each other: 2:1 to 3:1 ratio, attack around 10 to 20 ms so consonant attacks get through, release between 100 and 150 ms, threshold set for 3 to 5 dB of gain reduction on the fuller passages. The second, faster and set higher, only handles the peaks that still get through: 4:1 to 6:1, short attack, another 2 to 4 dB. The total stays in the 6 to 9 dB range, but it is spread out, so it is inaudible.

A gentle expander upstream can complete the setup, provided it is set as an expander and not as a gate: range limited to 6 or 8 dB, low threshold, long recovery. A noise gate that slams between sentences is more distracting than the noise it removes.

Stage 4 — de-essing comes after, never before

The FabFilter Pro-DS de-esser and its sibilance detection display
FabFilter

Sibilance is almost never a problem on the raw take. It becomes one afterwards, mechanically, because compression pulls down everything sustained without touching short transients, and because the presence lift lands just below the band where sibilance lives. That is why the de-esser belongs at the output of the compression stage, not at its input.

Three moves are enough. Identify the working band by soloing the detection path: 5 to 7 kHz on a low voice, 6 to 9 kHz on a bright one. You should hear “s” and “sh” sounds, not whole words. Then set the threshold for 3 to 6 dB of reduction on peaks only, with the meter resting most of the time. Finally, favour split-band operation over wideband: the first only attenuates the band concerned, the second ducks the entire voice on every sibilant and produces a very recognisable reverse-lisp effect.

Stage 5 — calibrate in LUFS, not in dBFS

This is where the gap between an amateur edit and a professional delivery shows most clearly. Normalising an export to -3 dBFS calibrates nothing: dBFS measures an instantaneous peak, not a sensation of level. Two files normalised to the same peak can differ by 8 dB by ear. The relevant measurement is LUFS, which integrates level over time with a weighting close to the ear’s own sensitivity.

Loudness targets by destination -24 LKFS — North American television (ATSC A/85) -23 LUFS — European television and radio (EBU R128) -19 LUFS — mono podcast reference -16 LUFS — stereo podcast -14 LUFS — YouTube and music platforms -29 -27 -25 -23 -21 -19 -17 -15 -13 -11 -9 Integrated level in LUFS. Further right means the programme is perceived as louder. In every case the true peak ceiling stays at -1 dBTP.
The same voice delivered at -23 LUFS for a broadcaster and at -14 LUFS for YouTube is not mixed differently: it is only calibrated differently on the way out.

Video platforms now normalise on playback. Delivering a file far louder than the target will not make it louder for the viewer: the player will pull it back down, and all that remains is the dynamic range crushed along the way. On YouTube, aiming for -14 LUFS with a true peak ceiling at -1 dBTP gives a clean result and avoids encoding distortion. For a podcast, the reference is -16 LUFS in stereo, with an equivalent around -19 LUFS when the programme is distributed in mono.

Television is another world, and the one with the strictest constraints. EBU R128, applied across Europe, sets an integrated level of -23 LUFS with a tolerance of plus or minus 0.5 LU on programmes, and a -1 dBTP ceiling. In the United States, ATSC A/85 settles on -24 LKFS. A piece delivered to a broadcaster without that calibration comes straight back, however good the edit. The detail of these standards, from integrated measurement to loudness range, is covered in our article on broadcast loudness.

The true peak headroom is not bureaucratic detail. A file touching 0 dBFS before encoding produces overshoots once compressed to AAC or Opus, and therefore crackles on the viewer’s player. dBTP measures exactly that reconstructed signal between samples, and that is the number to watch.

What is already sitting in your editing software

Final Cut Pro, whose audio inspector gathers voice isolation and equalisation
Apple

None of the three major editing applications forces you out to a dedicated audio workstation to treat a voice properly. They simply do not put the tools in the same place, and do not call them by the same names.

StageFinal Cut ProPremiere ProDaVinci Resolve
CleanupVoice Isolation and noise reduction in the audio inspectorEnhance Speech, DeNoise and DeReverb in the Essential Sound panelVoice Isolation, Noise Reduction and De-Hummer in the Fairlight FX
EqualisationGraphic and parametric equaliser in the inspectorParametric equaliser and dialogue presetsSix-band equaliser on every channel strip of the Fairlight page
DynamicsCompressor, limiter and expander in the audio effectsMultiband compressor, DeEsser and Essential Sound dynamicsCompressor, expander and gate built into every channel, plus the Dialogue Processor
Loudness meteringThe weakest of the three: use a metering plug-in or verify on exportLoudness Radar effect and amplitude statisticsBuilt-in Loudness Meter with EBU R128 and ATSC A/85 presets

On the single criterion of voice treatment, the Fairlight page in DaVinci Resolve remains the best equipped of the three: a real console strip per track, proper buses, a loudness meter that matches broadcast standards, and a processing engine designed around dialogue. Premiere Pro compensates with the ergonomics of its Essential Sound panel, which puts the right parameters within reach without switching pages. Final Cut Pro is the fastest for simple treatment and the most limited as soon as measurement is required.

Making voice and music share the same mix

The last reflex worth correcting concerns the music bed. The rule of thumb that circulates — sit the music 6 dB under the voice — consistently ends up too loud. On a speech-led programme, the useful gap sits between 12 and 18 dB while someone is talking. Music does not need to be audible during speech, it needs to be present.

Automatic ducking does the job provided the recovery time is long, in the region of 600 ms to one second. Set short, it makes the music breathe on every syllable and immediately draws attention to itself. A more elegant solution is to dynamically carve the music only in the voice’s presence band, around 2 to 3 kHz, with a dynamic equaliser triggered by the dialogue bus: the music keeps its low end and its top, and simply frees up the intelligibility band. How that cross-triggering works is detailed in our guide to sidechaining.

For productions headed to air, remember that one more processing stage will be applied after you, at transmission. A mix that already leaves the edit heavily compressed comes out squashed once it has been through the broadcast chain: better to deliver a calibrated and reasonably dynamic signal, as our article on on-air processing explains.

The six mistakes that come up most often

  1. Normalising the export to -3 dBFS and believing the programme is calibrated.
  2. Equalising before cleaning, and therefore equalising noise along with the voice.
  3. Handing the whole dynamic range to a single, hard-driven compressor.
  4. Placing the de-esser at the head of the chain, where it achieves nothing.
  5. Signing off the mix on headphones only, without ever checking on a mediocre loudspeaker.
  6. Delivering a duplicated mono voice in stereo with asymmetric processing on the two sides.

The walkthrough on video

Blackmagic Design offers a complete introduction to the Fairlight page, still the best entry point for understanding how a console strip works inside a video editing application.

None of the above requires buying anything. The six stages fit inside the modules shipped with the three dominant editing applications, and the whole chain takes about twenty minutes to set up on a typical project before it can be saved as a preset. The payoff is immediate: a voice that stays intelligible on a phone in a noisy street, a level that no longer jumps between sequences, and an export that passes compliance checks first time.

Share.

About Author

After 20+ years in professional audio: live sound engineering, studio technical direction (Deep Forest, Pierre Jacquot), head of digital marketing at Playback.fr. A first-hand witness to the analog-to-digital shift, I track the whole audio landscape and break it down here — no fluff.