The short version
- A clean recording is not a finished voice: without a processing chain, dialogue stays dull, uneven and fragile as soon as music comes in.
- The order of the stages matters as much as their settings — cleanup, high-pass and subtractive EQ, compression, de-essing, polish, then level calibration.
- The loudness target depends on the destination: around -14 LUFS on YouTube, -16 LUFS for a podcast, -23 LUFS for European television.
- Final Cut Pro, Premiere Pro and DaVinci Resolve can all do this work. What is missing is the method, not the tools.
The cut is locked, the grade is clean, the titles are in place — and the voice still sounds exactly like the file that came off the recorder. This has become so common that it now passes for normal: in most video production pipelines, dialogue travels all the way through without anyone really treating it. The recording was looked after, a decent microphone was placed at a sensible distance, levels were watched on set, and the audio job is assumed to be done.
It is not. A clean recording is raw material, not a result. Between the original file and a voice that holds its own against a music bed, in headphones as well as on a phone speaker, sits a processing chain of five or six stages. That chain is what is almost always missing, and this article walks through it stage by stage, tool by tool.
The problem is not the recording, it is the missing chain
A video editor works the image in successive layers: white balance, primaries, secondaries, masks, grain. Nobody would claim a shot is graded because the exposure was pushed half a stop. On audio, though, the gesture often stops there: pull the fader up until it can be heard, move to the next shot.
Spoken voice, however, is a signal with a very wide dynamic range. Between a syllable attacked at the start of a sentence and a breathed-out ending, 20 to 30 dB of difference is routine. No fader position fixes that: if the start of the sentence sits right, the end disappears under the music; if the end is audible, the start clips. A fader sets an average level, it does not shape a dynamic range.
The two symptoms that give away a missing chain are always the same. First, intelligibility collapses the moment a music bed arrives, because the voice was never cleared in the frequencies where music masks it. Second, perceived level jumps from one sequence to the next, because the takes come from different days, different rooms and sometimes different microphones, and no stage was ever tasked with bringing them to a common denominator.
The chain, in order
That sequence is not an aesthetic convention, it is a logical constraint. A noise reducer analyses a spectrum: equalise before it and you hand it a spectrum that has already been bent. A compressor follows an envelope: leave in the sub-bass rumble of a lavalier rubbing against a shirt and it will trigger on events nobody can hear. A de-esser looks for sibilance: place it upstream and it works on a signal where sibilance is not yet a problem, and lets it straight back in at the next stage.
Stage 1 — clean before you treat

Four nuisances turn up on almost every dialogue take shot in real conditions: broadband hiss from a preamp pushed hard on an insensitive microphone, continuous background noise from ventilation or traffic, mains hum at 50 or 60 Hz with its harmonics, and room reverberation. The first three respond well to treatment. The fourth is the hardest, because it is correlated with the voice itself.
Work from the steady to the transient. A noise reducer trained on a profile — two seconds of silence taken from the same recording, not from another one — does most of the job with 6 to 9 dB of reduction, and no more. Past that, the ear hears the processing before it hears the voice: word endings turn metallic and the room ambience drops out abruptly between lines. A de-hummer set to the mains fundamental then clears 50 or 60 Hz and their multiples. Mouth clicks and heavy breaths are dealt with last, one by one, by hand or with a dedicated module.
Neural voice isolators are not meant to run at 100 %
Voice Isolation in Final Cut Pro and DaVinci Resolve Studio, Enhance Speech in Premiere Pro: these neural isolators are spectacular on a thirty-second demo and destructive across a forty-minute interview. Pushed all the way, they build a vacuum-packed voice with no ambience at all, dropped sentence endings and metallic artefacts on fricatives. The right approach is to dial them between 30 and 60 % and let the room live, topping it up with a neutral ambience bed under the dialogue if needed. We covered this whole family of tools in our round-up of real-time voice isolation tools.
Stage 2 — EQ starts with what you remove
The most widespread reflex, and the most counterproductive, is to open an equaliser and lift the top end until the voice cuts through. It does cut through, along with the hiss, the sibilance and the mouth noise. The right method is the opposite: take out what is in the way first, then lift far less than you expected to.
Set the high-pass first, with a steep slope, between 70 and 90 Hz on a male voice and between 100 and 120 Hz on a female one. The reliable method is to raise the corner frequency until the voice starts to sound thin, then back off one notch. Anything below the fundamental adds nothing to the timbre and makes the following stages work for nothing.
Next comes the 200 to 400 Hz band, responsible for the cardboard-box effect. It is particularly loaded on lavaliers worn under clothing and on close-miked takes, where proximity effect swells the low mids. A wide bell with 2 to 4 dB of cut is almost always enough. Between 400 Hz and 1 kHz live nasality and room colouration: there you need a narrow bell, swept until you find the exact spot, and a surgical cut of 3 to 5 dB.
Only after that cleanup do you lift presence, between 2.5 and 5 kHz, with a wide bell and 2 to 3 dB at the very most. The air shelf above 10 kHz comes last, gently, and never on a noisy take — it would bring the hiss up with everything else. The behaviour of each filter type is broken down in our guide to equalisation.
Stage 3 — two gentle compressions beat one brutal one

Compression is the stage video editors get wrong most often, because they treat it as a level corrector rather than an envelope shaper. A single compressor asked to absorb 10 to 12 dB of dynamic range produces an instantly recognisable result: the voice pumps, breaths jump to the front, and the room floor rises and falls with every sentence.
The fix is to split the work across two stages. The first one, slow, evens out sentences against each other: 2:1 to 3:1 ratio, attack around 10 to 20 ms so consonant attacks get through, release between 100 and 150 ms, threshold set for 3 to 5 dB of gain reduction on the fuller passages. The second, faster and set higher, only handles the peaks that still get through: 4:1 to 6:1, short attack, another 2 to 4 dB. The total stays in the 6 to 9 dB range, but it is spread out, so it is inaudible.
A gentle expander upstream can complete the setup, provided it is set as an expander and not as a gate: range limited to 6 or 8 dB, low threshold, long recovery. A noise gate that slams between sentences is more distracting than the noise it removes.
Stage 4 — de-essing comes after, never before

Sibilance is almost never a problem on the raw take. It becomes one afterwards, mechanically, because compression pulls down everything sustained without touching short transients, and because the presence lift lands just below the band where sibilance lives. That is why the de-esser belongs at the output of the compression stage, not at its input.
Three moves are enough. Identify the working band by soloing the detection path: 5 to 7 kHz on a low voice, 6 to 9 kHz on a bright one. You should hear “s” and “sh” sounds, not whole words. Then set the threshold for 3 to 6 dB of reduction on peaks only, with the meter resting most of the time. Finally, favour split-band operation over wideband: the first only attenuates the band concerned, the second ducks the entire voice on every sibilant and produces a very recognisable reverse-lisp effect.
Stage 5 — calibrate in LUFS, not in dBFS
This is where the gap between an amateur edit and a professional delivery shows most clearly. Normalising an export to -3 dBFS calibrates nothing: dBFS measures an instantaneous peak, not a sensation of level. Two files normalised to the same peak can differ by 8 dB by ear. The relevant measurement is LUFS, which integrates level over time with a weighting close to the ear’s own sensitivity.
Video platforms now normalise on playback. Delivering a file far louder than the target will not make it louder for the viewer: the player will pull it back down, and all that remains is the dynamic range crushed along the way. On YouTube, aiming for -14 LUFS with a true peak ceiling at -1 dBTP gives a clean result and avoids encoding distortion. For a podcast, the reference is -16 LUFS in stereo, with an equivalent around -19 LUFS when the programme is distributed in mono.
Television is another world, and the one with the strictest constraints. EBU R128, applied across Europe, sets an integrated level of -23 LUFS with a tolerance of plus or minus 0.5 LU on programmes, and a -1 dBTP ceiling. In the United States, ATSC A/85 settles on -24 LKFS. A piece delivered to a broadcaster without that calibration comes straight back, however good the edit. The detail of these standards, from integrated measurement to loudness range, is covered in our article on broadcast loudness.
The true peak headroom is not bureaucratic detail. A file touching 0 dBFS before encoding produces overshoots once compressed to AAC or Opus, and therefore crackles on the viewer’s player. dBTP measures exactly that reconstructed signal between samples, and that is the number to watch.
What is already sitting in your editing software

None of the three major editing applications forces you out to a dedicated audio workstation to treat a voice properly. They simply do not put the tools in the same place, and do not call them by the same names.
| Stage | Final Cut Pro | Premiere Pro | DaVinci Resolve |
|---|---|---|---|
| Cleanup | Voice Isolation and noise reduction in the audio inspector | Enhance Speech, DeNoise and DeReverb in the Essential Sound panel | Voice Isolation, Noise Reduction and De-Hummer in the Fairlight FX |
| Equalisation | Graphic and parametric equaliser in the inspector | Parametric equaliser and dialogue presets | Six-band equaliser on every channel strip of the Fairlight page |
| Dynamics | Compressor, limiter and expander in the audio effects | Multiband compressor, DeEsser and Essential Sound dynamics | Compressor, expander and gate built into every channel, plus the Dialogue Processor |
| Loudness metering | The weakest of the three: use a metering plug-in or verify on export | Loudness Radar effect and amplitude statistics | Built-in Loudness Meter with EBU R128 and ATSC A/85 presets |
On the single criterion of voice treatment, the Fairlight page in DaVinci Resolve remains the best equipped of the three: a real console strip per track, proper buses, a loudness meter that matches broadcast standards, and a processing engine designed around dialogue. Premiere Pro compensates with the ergonomics of its Essential Sound panel, which puts the right parameters within reach without switching pages. Final Cut Pro is the fastest for simple treatment and the most limited as soon as measurement is required.
Making voice and music share the same mix
The last reflex worth correcting concerns the music bed. The rule of thumb that circulates — sit the music 6 dB under the voice — consistently ends up too loud. On a speech-led programme, the useful gap sits between 12 and 18 dB while someone is talking. Music does not need to be audible during speech, it needs to be present.
Automatic ducking does the job provided the recovery time is long, in the region of 600 ms to one second. Set short, it makes the music breathe on every syllable and immediately draws attention to itself. A more elegant solution is to dynamically carve the music only in the voice’s presence band, around 2 to 3 kHz, with a dynamic equaliser triggered by the dialogue bus: the music keeps its low end and its top, and simply frees up the intelligibility band. How that cross-triggering works is detailed in our guide to sidechaining.
For productions headed to air, remember that one more processing stage will be applied after you, at transmission. A mix that already leaves the edit heavily compressed comes out squashed once it has been through the broadcast chain: better to deliver a calibrated and reasonably dynamic signal, as our article on on-air processing explains.
The six mistakes that come up most often
- Normalising the export to -3 dBFS and believing the programme is calibrated.
- Equalising before cleaning, and therefore equalising noise along with the voice.
- Handing the whole dynamic range to a single, hard-driven compressor.
- Placing the de-esser at the head of the chain, where it achieves nothing.
- Signing off the mix on headphones only, without ever checking on a mediocre loudspeaker.
- Delivering a duplicated mono voice in stereo with asymmetric processing on the two sides.
The walkthrough on video
Blackmagic Design offers a complete introduction to the Fairlight page, still the best entry point for understanding how a console strip works inside a video editing application.
None of the above requires buying anything. The six stages fit inside the modules shipped with the three dominant editing applications, and the whole chain takes about twenty minutes to set up on a typical project before it can be saved as a preset. The payoff is immediate: a voice that stays intelligible on a phone in a noisy street, a level that no longer jumps between sequences, and an export that passes compliance checks first time.