Explainer bench — audio

Same 10.5-second beat, same picture, same narration take. Only the audio changes. Every number on this page is measured off the actual render with ffmpeg, not estimated.
Visuals (#2259) Audio (#2260)

1 · What was actually wrong

Measured on the six videos already shipped, before anything was changed.
Shipped videoIntegratedRangeTrue peak
atomic-habits-60s-24.3 LUFS2.7 LU-4.4 dBTP
atomic-habits-v2-22.9 LUFS1.4 LU-7.0 dBTP
atomic-habits-v3-21.4 LUFS4.5 LU-4.9 dBTP
atomic-habits-v4-24.4 LUFS2.1 LU-6.6 dBTP
chimp-paradox-30s-24.0 LUFS2.1 LU-6.9 dBTP
rich-dad-30s-24.2 LUFS1.7 LU-7.4 dBTP

The raw ElevenLabs narration measures -24.3 LUFS and the finished mixes measure -21.4 to -24.4. Nothing was ever loudness-normalised — whatever the voice happened to come back at is what shipped. YouTube normalises downward only: it turns loud uploads down and never turns a quiet one up (verified — its reference is widely reported as about -14 LUFS, though YouTube publishes no official figure). So these play roughly 10 dB below every video around them. You turn the phone up to compensate, and every transition comes up with it.

The three SFX files the v3 build used are still on disk. Their true peaks measure -13.9, -21.5 and -25.5 dBTP — an 11.6 dB spread — and the build script then multiplied them by hand-picked constants (0.62 on the whoosh, 0.4 on the chime). So the whoosh landed 7.6 dB above the chime for no reason other than which file it was. pop.wav is also genuinely clipped (21 samples pinned at full scale). Across the finished v3 mix the loudest transitions sit 5.5 LU above the narration.

Honest verdict: the dominant defect was levels and the total absence of a target, not the sounds themselves. The old whoosh is a perfectly usable sound that was placed by guesswork. Only pop.wav was actually broken.

2 · Old mix vs the new spec

Same picture, same narration, same SFX placements. Only the levelling rules differ.

3 · Intro theme — four candidates

2 to 3.2 seconds each. Tap to hear the theme on its own, or play the full clip underneath to hear it lead into the video.

4 · The SFX pack — 30 sounds

Tap any of them. Every file is synthesised in the repo (so CC0, no attribution and no licence to track) and normalised to the same measured level — the spread across the whole pack is 1.99 LU, against 11.6 dB for the three files it replaces.

5 · The mix spec

Written down and enforced in code, not by ear. Every level is an offset from one anchor — the narration.
ElementRuleWhy
Final mix-14.0 LUFS, -1.0 dBTPthe platform reference; anything quieter just plays quieter
Narration-16.0 LUFS, anchorevery other level is an offset from this, so it survives a voice change
Music bed-14 LU under the voiceaudible in a gap, inaudible over a word
Music ducking-8 dB, 40 ms / 420 msreal sidechain keyed off a dB-domain voice envelope
SFX bus-9 LU under the voiceone number that means the same thing for every sound in the pack
SFX ducking-3 dBso a transition mid-sentence cannot mask a word
Per categoryshake 0 → counter -8 dBa camera shake is meant to be felt; a counter tick fires twenty times
Hard rule≤ 1 LU over the voicechecked on every render; the build fails if it breaks