Bonus · A deeper dive into Figure 3.21

Sirens Are FM Running Slow

Wail, yelp, and phaser aren’t different machines. They’re one carrier under frequency modulation at three different rates. Raise that one rate past about 20 Hz and the sweep stops being a gesture in time and becomes a spectrum: sidebands, timbre, an instrument.

Builds on Chapter 3 · Spectra, Modulation & Phase — frequency modulation.

Chapter 3’s Live FM demo (Figure 3.21) builds FM tones one preset at a time. This page is its deeper-dive companion, the way the Fourier Sandbox extends the harmonics figure: the presets become nine points on a continuous plane, and the plane makes the single idea of the chapter audible in one drag. A siren is the one FM signal already heard thousands of times, so the slow end starts from recognition rather than from an abstraction.

The trick is that vibrato, a siren, and a bell all come from the same patch: an oscillator whose frequency is wobbled by a second one. Hold the depth of that wobble constant and change only its rate, and a perceptual border is crossed. Below the transition an event is heard; above it, the ear stops tracking the excursion and integrates it into a steady tone color. One number changes; the category of the percept changes.

FIG. 3·S

The FM Plane

Press Start audio, then drag anywhere on the plane below (or use the sliders). Left–right is the modulation rate fm; up–down is the peak deviation d. Because the modulation index is I = d/fm, the faint 45° diagonals are lines of constant I — the index is a place on the plane, not just a number in a readout. The dashed vertical rule at fm = 20 Hz is a conventional landmark rather than a hard boundary: the change from hearing a rhythm to hearing a timbre happens gradually, over a transition region somewhere around 8–20 Hz that depends on the sound and on the listener. Then press Run the collapse and watch the cursor cross it.

audio suspended
Spectrogram · 0–4 kHz · linear · time →
Spectrum · linear Hz · carrier sideband stems = Bessel prediction
Mod. index I = d / fm
Bandwidth (Carson)
Ratio fm : fc
Regime
The FM plane · drag to reposition · 45° rules = constant modulation index fm → · d ↑

Carrier

Modulator

Motion & analysis

Presets, ordered by modulation rate — press 17fm

Reading the panels. The spectrogram uses a linear frequency axis (unlike the log-frequency spectrograms elsewhere in Chapter 3): FM sidebands sit at even spacing fm apart, so the comb only reads as a comb on a linear axis. The spectrum’s colored stems are a prediction, not a fit: their heights are the Bessel amplitudes |Jn(I)|, shape-matched to the measured peak. Watch the carrier stem (red) drop into the floor near I ≈ 2.405 — the first zero of J0, where all the energy has fled to the sidebands. The overlay is suppressed below fm ≈ 15 Hz, where the sideband picture is meaningless.

Face control · optionalCamera-captured imagery is neither saved nor transmitted.
Jaw open → deeper sweep
Brows raised → carrier up to an octave higher
Lips pursed → rate up to ten times faster
Each gesture is a departure from the current setting rather than a replacement for it, so a resting face sounds exactly like whichever preset or slider position it started from, and every preset keeps its own character. Mapping jaw aperture to the modulation index isn’t arbitrary — it is roughly what the mouth does when imitating a siren. A face-tracking model runs entirely in the browser; camera-captured imagery is neither saved nor transmitted.

Why the border sits near 20 Hz

Below roughly 20 Hz the ear follows the frequency excursion as it happens: the pitch is heard to rise and fall, a sweep or a warble in time. Above it, the excursion repeats faster than the ear can track, and the auditory system does what it always does with anything periodic — it hears a pitch and a color. The zig-zag in the spectrogram freezes into a static comb of sidebands, each one spaced fm from its neighbour. Nothing in the patch changed but one number. This is the same border that separates the slow LFO of vibrato from the audio-rate modulation that builds a timbre, and it is the whole reason FM synthesis works.

Doppler, honestly

Switch on the Doppler pass-by and the source crosses in front of the listener. The pitch shift is computed from the radial velocity and applied to both the center frequency and the deviation — the entire spectrum scales, as it physically must, not just the middle of it. (Most in-browser Doppler demos get this wrong.) The spatialization that puts the siren out in front of the listener, and lets it swing from one side to the other, is the same head-related rendering explored in Chapter 9’s auralization material.

💡

Try this: set the deviation with the XY pad (the FM plane above) about two-thirds up, then drag straight rightward along a single horizontal line — holding d fixed while only fm climbs. That horizontal drag is the collapse, under a single finger: a wail, a warble, a rough buzz, and then, as the dashed rule is crossed, a pitched instrument.

Sources & Inspiration

  • FM synthesis. John M. Chowning, “The Synthesis of Complex Audio Spectra by Means of Frequency Modulation,” J. Audio Eng. Soc. (1973) — the discovery that modulating a carrier’s frequency generates rich sidebands, so a slow siren sweep and a bright timbre are one process at different rates. Commercialized in Yamaha’s DX7.