跳转到主要内容
Workflows 5 min read

Mixing Basics: Layering Voice, Music and Effects by Volume

When multiple audio tracks become one, the classic failure is music drowning the voice and levels jumping around. This guide covers mixing clean, layered audio with local tools — enough for podcast beds, voiceover backing and simple dubbing.

Think in layers: one main track, everything else supports it

Decide first which track is primary. In a talking-head video it is the voice; in a travel vlog it may be ambient sound or narration. Set the main track’s level as the reference and bring everything else down relative to it — this decision is eighty percent of mixing.

A common starting point for music beds is 10–15 dB below perceived voice level, which in volume terms means multiplying the music track by roughly 0.3–0.4. Do not tune by feel in one pass: start from this ratio, listen, then fine-tune against the content.

Effect tracks (transitions, stings) follow the rule “noticeable at the moment they play, invisible otherwise”, typically louder than the bed but below the voice. Assigning every track a role before touching any knob matters more than any parameter.

Normalize first: align loudness per track before mixing

Loudness varies wildly between sources: phone-recorded voice may be quiet, downloaded music may be loud. Normalize each track to a similar loudness level before mixing — otherwise the volume ratios you set are meaningless.

Only then do relative adjustments: main track untouched, bed multiplied by 0.3, effects by 0.6, and listen. This order — normalize, then layer — cuts rework roughly in half compared with mixing raw tracks.

Processing tracks separately has a second benefit: problems (clipping, hiss) become obvious per track instead of masking each other. Layered listening is the fundamental debugging skill of mixing.

Assemble and align: combining tracks into one file

With levels set, use the merge tool to combine music and voice into one file. Align lengths first: trim a longer bed, loop-pad a shorter one — never force a mismatch.

Timing design affects listening more than volume does: let the bed enter two or three seconds before the voice, keep a few seconds of fade at the end after the last word. The audio fade tool builds this breathing room at almost zero cost.

For sequenced structures (intro bed → voice → interlude → outro), assemble segment by segment with the merge tool. Local tools favor this pattern: process pieces separately, then join in order.

Classic failure modes: hiss, pops and doubled voices

If the mix sounds muddy, the usual culprit is layered noise floor: two tracks each with faint hiss sum into audible noise. De-noise the offending track before mixing; de-noising after the fact is a rescue measure, not correct procedure.

Pops happen at splice points where the waveform jumps. Add very short fades (tens of milliseconds) to the head and tail of every segment and splice pops essentially vanish.

A hollow, underwater-sounding voice usually means two similar tracks stacked — for example live voice plus a later dub of the same content. When the same content exists twice, keep one and delete the other; this beats any parameter tweak.

Export choices: format and loudness by destination

Export by use: AAC inside video for platforms; MP3 for standalone audio release (compatibility) or FLAC (archival). Keep intermediate work files in lossless WAV/PCM so repeated processing never compounds lossy artifacts.

Do not slam the final level: leave some peak headroom so platform-side loudness normalization does not squash the mix. Loudness consistency plus healthy dynamics is what “loud” should actually mean.

Mixing is iterative: export, then listen on another device — phone speaker, earbuds, desktop speakers — and fix whatever is off. Three rounds usually stabilizes a mix; do not over-invest in monitoring gear.

Frequently asked questions