Key points
- In a podcast the voice is not one element of the mix, it is the mix. Everything else is decoration.
- Mic distance and room acoustics decide the result before a single plugin is loaded.
- One track per speaker, never a recorded blend: without that, no level matching is possible at the edit stage.
- Delivery target: -16 LUFS in stereo, around -19 LUFS in mono, with a true-peak ceiling at -1 dBTP.
A podcast has no picture, no set and no cutting rhythm to rescue a mediocre voice. Listeners are on thirty-euro earbuds, in a car, or in a kitchen with the extractor fan running, and they decide within about thirty seconds whether they stay. What is at stake in those thirty seconds is not the topic of the episode: it is listening comfort.
The good news is that spoken-word production is technically simple. There are no instruments, no complex imaging, no arbitration between twenty sources. The trade-off is that nothing hides: a hiss, a live room, a guest twice as loud as the host, and the flaw fills the whole picture. Here is the complete chain, from mic placement to export.
It all starts with distance to the microphone

The condenser-versus-dynamic debate takes up a lot of space in podcasting discussions, and it is largely secondary. What decides the result is the ratio between direct sound and room sound, and that ratio depends first on distance.
At 40 cm from a microphone, in an ordinary room, the reverberant share becomes audible and distracting. At 10 or 15 cm it drops into the background. That single difference matters more than the choice of capsule. It also explains why a cardioid dynamic, less sensitive and worked up close, often beats a better condenser in an untreated room: it hears less of the room because you speak closer to it.
Three practical rules follow. Sit about a fist’s width from the grille, slightly off axis so plosives pass beside the capsule rather than into it. Add a pop filter, which fixes the working distance as much as it tames the p’s. And mount the mic on a boom arm rather than standing it on the desk, otherwise every keystroke and every mug being moved travels up through the structure.
The room matters more than the microphone
A podcast studio does not need to look good, it needs to sound dead. The absolute priority is the wall behind the speaker, because that is the surface sending the voice straight into the back of the microphone. Then the opposite wall, then the desktop itself, which produces a very short reflection responsible for the comb-filtered colouration that makes a voice sound hollow for no obvious reason.
Two absorbers of 60 by 120 cm, 10 cm thick, correctly placed, transform an ordinary room. A rug, heavy curtains and a full bookcase do part of the job for very little money. What does not work is thin foam tiles glued up at random: they absorb the highs and leave the low mids untouched, producing a room that is dull and still muddy. The principles of absorption, diffusion and modal control are covered in our guide to studio acoustic treatment.
One track per speaker, always

This is the decision that governs everything downstream. Recording a stereo blend of the conversation saves three clicks and permanently rules out fixing anything: no rescuing the guest who drifted off mic, no removing spill from the neighbouring microphone, no matching levels. Every current podcast console records multitrack to SD card or over USB, and that capability should be used systematically.
Remote guests deserve separate treatment. A videoconference stream is compressed, bandwidth-limited and at the mercy of the network: it is not usable material for a careful production. The solution is to have each participant record locally on their own device, and to use the real-time stream only as a synchronisation reference. The feed sent back to that guest must be a mix-minus, meaning the whole programme except their own voice, otherwise they hear themselves echoed back and stop speaking naturally. The principle and its wiring are explained in our article on mix-minus.
Edit before you mix

Editing a spoken-word programme is not mixing work, it is editorial work. It means removing hesitations, repetitions, false starts and over-long silences, then joining everything back inaudibly. Voice-oriented workstations build their interface around that task, with automatic region-based levelling and cuts that preserve ambience continuity.
Three traps recur. The first is removing every silence: dialogue without breathing becomes suffocating to listen to and unreadable over an hour. Remove the dead time, not the punctuation. The second is cutting in the middle of a breath, which creates an artefact the ear catches instantly — cut on silence, or crossfade over a few milliseconds. The third is editing by looking at the waveform instead of listening: a 400 ms gap can be perfectly natural or plainly awkward depending on what precedes it.
On multi-mic sessions recorded in the same room, spill is the real subject. Every microphone also picks up the other speakers, slightly delayed and with degraded timbre. Muting tracks outside their own speech — by hand, or with a gentle expander — immediately tightens the image and removes the sense of the room rising whenever several mics are open.
The processing chain, adapted to spoken word

The sequence is the same as for a video edit — clean-up, high-pass, subtractive EQ, compression, de-essing, loudness calibration — and we covered it stage by stage in our article on the voice processing chain in a video edit. Three podcast-specific points are worth underlining.
First, compression can be more assertive than in drama or documentary. Podcasts are mostly consumed on the move, in noisy environments, on poor transducers. A residual dynamic range of 8 to 10 dB between loudest and quietest works well; beyond that, listeners spend their time on the volume control. Two stages of 4 to 5 dB each still beat a single one pushed hard.
Second, processing happens track by track, before any bus processing. Every voice has its own fundamental, its own sibilance region and its own dynamics: a single setting applied to the blend works for one speaker and damages the other. The master bus carries only a transparent limiter and the final calibration.
Third, clean-up should stay light. A podcast recorded in a decent room does not need a neural isolator pushed to the limit; it needs 4 to 6 dB of noise reduction and a well-set high-pass. Voice separation tools, of which we published a full overview, belong to genuinely compromised recordings.
| Setting | Low voice | Mid voice | Bright voice |
|---|---|---|---|
| High-pass | 70 to 80 Hz | 85 to 100 Hz | 100 to 120 Hz |
| Low-mid cut | 250 Hz, -3 dB | 300 Hz, -2 dB | 350 Hz, -2 dB |
| Presence lift | 3.5 kHz, +2 dB | 3 kHz, +2 dB | 2.5 kHz, +1.5 dB |
| De-esser band | 5 to 7 kHz | 6 to 8 kHz | 6.5 to 9 kHz |
| Total gain reduction | 7 to 9 dB | 7 to 9 dB | 6 to 8 dB |
Matching speakers without flattening them
The level gap between a host used to the microphone and a guest talking into thin air is the most common flaw in interview podcasts, and the easiest to fix. The goal is not to make everyone sound the same — timbres must stay distinct, that is what lets a listener follow a three-way conversation — but to reach equivalent perceived levels.
The method is mechanical. Isolate each participant’s contributions, measure the integrated level of each separately, and adjust track gain to bring everyone within one unit of each other. This is done before compression, not after: a compressor fed by sources at very different levels produces incomparable gain reduction from one track to the next, and the gap reappears in another form.
Calibrating and exporting
The reference target for a podcast delivery is -16 LUFS in stereo, with a true-peak ceiling at -1 dBTP. When the programme is distributed in mono, the equivalent landmark sits around -19 LUFS, the difference coming from the way level is integrated over a single channel. The main platforms normalise on playback, so delivering louder brings no benefit and costs dynamic range.
The mono-or-stereo question settles quickly. An interview podcast recorded on fixed microphones is a mono programme: delivering it in stereo doubles the bitrate for nothing and complicates life for players that sum the channels. Stereo earns its place when there is crafted music, atmospheres, or a deliberate spatial intent. When in doubt, mono is the right call, and it stays perfectly compatible with every platform.
On formats, MP3 at 128 kbps in mono or 192 kbps in stereo remains the safest compatibility compromise. The -1 dBTP ceiling makes full sense here: lossy encoding is precisely what manufactures overshoots on peaks, and therefore crackles on the listener’s end. These measurements are detailed in our loudness guide.
The gear that genuinely changes something
The hierarchy of useful spending is counter-intuitive. Acoustic treatment comes first, far ahead of everything else, because it fixes the one flaw no software recovers. Then the boom arm and pop filter, which lock in a constant working distance. The microphone itself only comes third: a decent cardioid dynamic used close in a treated room beats a high-end condenser placed fifty centimetres away in an empty living room.
Integrated podcast consoles do have real value, but rarely the one attributed to them. Their contribution is not sonic, it is organisational. They provide high-gain preamps suited to low-sensitivity dynamic mics, computer-free multitrack recording, independent headphone feeds for each participant and native mix-minus handling for phone callers. For field work, compact wireless systems now cover the same need in a pocket format, such as the Hollyland MELO P1.
The mistakes that cost the most
- Recording a blend instead of separate tracks: no correction remains possible.
- Treating the room last, after buying the microphone.
- Using the videoconference stream as final material for a remote guest.
- Applying one processing chain on the bus instead of treating each voice separately.
- Removing every breath and every silence at the edit.
- Exporting a fully mono programme in stereo, and above the target level.
The video walkthrough
RØDE walks through the voice processing built into its podcast console, a useful look at what these machines actually automate at capture time.
Nothing in this chain is a technical feat. What it demands above all is consistency: the same microphone, at the same distance, in the same room, with the same settings saved as a preset and the same target level at export. That regularity from one episode to the next, far more than sophisticated processing, is what makes a podcast listenable without fatigue over the long run.