Reference

Glossary

Key terms appearing in red throughout the chapters are collected here alphabetically. Each entry links back to the chapter where the term is first introduced and defined in context.

A

A-D
See analog-digital converter.Introduced in Chapter 4 — Sound Authoring & Casting.
A-weighting A-Bewertung
A frequency weighting applied to SPL measurements that emphasizes the mid-range where the ear is most sensitive (roughly the inverse of the 40-phon equal-loudness contour) and de-emphasizes the very low and very high frequencies. Reported as dB(A); the de facto standard for environmental and occupational noise readings. C-weighting and Z-weighting are flatter alternatives used for high-level or unweighted measurements.Defined here; not discussed in the chapter text.
absolute phase Absolutphase
The phase of a sinusoidal partial measured against an arbitrary time origin. For a steady tone the ear is largely insensitive to absolute phase, although it becomes audible when partials are mixed or when transient onsets are aligned.Introduced in Chapter 3 — Spectra, Modulation & Phase.
absorption coefficient Absorptionsgrad (Schluckgrad)
The fraction of sound energy a surface absorbs rather than reflects, written α and running from 0 for a perfect reflector to 1 for a perfect absorber. Tile and glass sit near the bottom; thick carpet, heavy drape and upholstered seating near the top. It is strongly frequency-dependent — soft furnishing takes treble far more readily than bass — which is why a furnished room sounds duller rather than merely quieter. Summed over a room it gives the room constant R = S̄α / (1 − ᾱ), which fixes the level of the reverberant field and hence the reverberation time.Introduced in Chapter 7 — Utilities & Effects.
acoustical computer-aided design (CAD) akustische CAD (computergestützte Akustikplanung)
Software used to model the acoustic behavior of a room or sound system before it is built. By combining geometric ray-tracing or image-source methods with auralization, an acoustical CAD package lets designers and clients audition a virtual space and compare alternative treatments.Introduced in Chapter 9 — 3D Sound & Auralization.
ADSR ADSR-Hüllkurve
An idealised four-segment amplitude envelope standard in synthesizers: Attack rises from zero to peak when a key is pressed, Decay falls to the Sustain level while the key is held, and Release falls back to zero when the key lifts. Real-instrument envelopes are richer but ADSR captures the gross gestures cheaply.Introduced in Chapter 3 — Spectra, Modulation & Phase.
aliasing Aliasing (Aliasfrequenzen / Spiegelfrequenzen)
Spurious low-frequency content produced when a signal contains frequencies above half the sampling rate (the Nyquist frequency) and is digitized without proper anti-alias filtering. The aliased frequencies fold down into the audible band as inharmonic tones that were not present in the original. Cured by band-limiting before the converter; see Nyquist theorem.Introduced in Chapter 6 — Digitization & Editing.
amplitude envelope Amplitudenhüllkurve
The slow-varying outline traced by the peak amplitude of a waveform over time, usually divided into attack, decay, sustain, and release segments. The envelope conveys much of an instrument’s identity, since the ear is highly sensitive to onset shape and decay rate.Introduced in Chapter 3 — Spectra, Modulation & Phase.
amplitude modulation Amplitudenmodulation (AM)
A process in which the instantaneous amplitude of a carrier signal is varied in proportion to a modulator signal. At sub-audio modulation rates this is heard as tremolo; at audio rates it produces sum-and-difference sideband partials around the carrier frequency.Introduced in Chapter 3 — Spectra, Modulation & Phase.
analog signal Analogsignal
A signal whose value varies continuously over time — for example, air pressure in sound, or the voltage on an audio cable. Analog signals carry sound between microphones, preamps, and loudspeakers before any digitization stage.Introduced in Chapter 1 — Communication, Frequency & Pitch.
analog-digital converter Analog-Digital-Wandler (A/D-Wandler)
A circuit that samples a continuous analog voltage at a fixed rate and assigns each sample an integer code, producing a stream of numbers suitable for digital storage and processing. Abbreviated A-D or ADC; the inverse operation is performed by a D-A.Introduced in Chapter 1 — Communication, Frequency & Pitch.
analysis-synthesis Analyse-Synthese
A two-stage processing paradigm in which a sound is first decomposed into a parametric description (such as a short-time spectrum or sinusoidal partials) and then resynthesized from that description, often with the parameters modified to alter pitch, duration or timbre. Phase vocoding and source-filter speech coding are classic examples.Introduced in Chapter 7 — Effects & DSP.
anechoic reflexionsfrei (schalltot)
Free of acoustic reflections. An anechoic chamber is a room whose surfaces absorb essentially all incident sound, and an anechoic recording captures only the direct sound of a source, making it suitable as a dry signal for later spatial or reverberation processing.Introduced in Chapter 9 — 3D Sound & Auralization.
audience Publikum (Zuhörerschaft)
The listeners for whom a production is intended. Awareness of audience — their age, cultural reference frame, listening environment and playback equipment — shapes nearly every authorial decision in audio.Introduced in Chapter 1 — Communication, Frequency & Pitch.
audio cast member Audio-Besetzungselement (Klangrolle)
Any sound — speech, music or effect — that has been selected and assigned a specific role within an audio production, by analogy with an actor cast in a film.Introduced in Chapter 4 — Sound Authoring & Casting.
audio interface Audio-Interface (Audioschnittstelle)
The hardware that mediates between analog audio signals and the computer, typically housing microphone preamps, A-D and D-A converters, headphone amplifiers and a digital connection (USB, Thunderbolt, etc.) to the host.Introduced in Chapter 1 — Communication, Frequency & Pitch.
audio mixing console Mischpult (Tonmischpult)
A device that combines multiple input signals into one or more output busses, providing per-channel gain, equalization, panning, auxiliary sends, and metering. Hardware consoles are still preferred for live work; software mixers serve the same role inside a DAW.Introduced in Chapter 5 — Miking & Recording.
auditory icon auditives Icon
An interface sound that signifies by resembling what it stands for — paper crumpling for a discarded file, a shutter for a photograph, a latch for something closing. Nothing needs to be learned, since the listener already knows what the world sounds like, but the vocabulary runs out quickly: most abstractions have no characteristic sound. Contrast earcon.Introduced in Chapter 4 — Sound Authoring & Casting.
auralization Auralisation
Rendering an acoustic environment so that a listener can hear how a source would sound in that space, by convolving a dry source with simulated room impulse responses and binaural filters. The audio counterpart of visualization in computer graphics.Introduced in Chapter 7 — Effects & DSP.
author Autor (Urheber)
The originator of a media work, who chooses its content and structure and thereby shapes the listener’s experience. In Sonic the term emphasizes intentional design of the audio chain from source to receiver.Introduced in Chapter 1 — Communication, Frequency & Pitch.
autocorrelation Autokorrelation
A measure of how strongly a signal resembles a delayed copy of itself, plotted against the delay (the lag). A periodic signal correlates with itself at every whole multiple of its period, so the lag of the first strong peak estimates the period — and its reciprocal the pitch. The measure reads periodicity from the waveform rather than from the spectrum, so it still reports the right rate when the fundamental partial is missing altogether, which is why it is one of the standard accounts of fundamental tracking and of virtual pitch. The same operation underlies pitch tracking in tuners, pitch correctors, and speech coders.Introduced in Chapter 3 — Spectra, Modulation & Phase.

B

back plate Gegenelektrode (Rückplatte)
The fixed, electrically charged electrode behind the diaphragm of a condenser microphone. The diaphragm and back plate form a capacitor whose capacitance — and therefore voltage — varies with sound pressure.Introduced in Chapter 5 — Miking & Recording.
balanced audio symmetrische Signalführung
A signal-cable scheme that carries the same audio on two wires of opposite polarity inside a shielded pair, with a third conductor as ground. Common-mode noise picked up equally on both signal wires is rejected by the differential input at the receiving end. XLR cables and TRS connectors used at line level are typical balanced interconnects; consumer RCA and TS cables are unbalanced (single signal wire + ground) and are more susceptible to hum over long runs.Defined here; not discussed in the chapter text.
band-pass (BPF) Bandpass(filter)
A filter that passes a contiguous band of frequencies between a lower and an upper cut-off and attenuates frequencies outside that band. Specified by center frequency and bandwidth (or Q).Introduced in Chapter 7 — Effects & DSP.
band-stop (BSF) Bandsperre (Notchfilter)
A filter that attenuates a contiguous band of frequencies while passing those above and below it; also called a band-reject or notch filter. Narrow band-stops are used to remove hum or other steady interference.Introduced in Chapter 7 — Effects & DSP.
beat frequencies Schwebungsfrequenz (Schwebungen)
Slow amplitude pulsations heard when two tones of nearly equal frequency sound together; the beat rate equals the difference between the two frequencies. Beats vanish as the tones are tuned to unison and are the basis for tuning by ear.Introduced in Chapter 1 — Communication, Frequency & Pitch.
BGM Hintergrundmusik (BGM)
Background music: the underscore that runs beneath dialogue, narration or visuals to set mood and pacing without drawing primary attention. The abbreviation is common in Japanese broadcast and post-production practice.Introduced in Chapter 1 — Communication, Frequency & Pitch.
bi-directional Acht (Achtercharakteristik / bidirektional)
A microphone directivity pattern that is equally sensitive to sound arriving from the front and rear of the diaphragm but rejects sound from the sides; also known as figure-8. Native to pressure-gradient ribbon elements and used for stereo mid-side and Blumlein arrays.Introduced in Chapter 5 — Miking & Recording.
binaural binaural
Pertaining to both ears. Binaural recording or rendering preserves the interaural time, level and spectral cues that the head and pinnae normally impose, so that headphone playback can convey a convincing externalized 3D image.Introduced in Chapter 9 — 3D Sound & Auralization.
bit depth Bittiefe (Quantisierungsauflösung)
The number of bits used to encode each sample in a PCM digital audio stream. Higher bit depth lowers quantization noise and widens dynamic range — roughly 6 dB per bit (see SQNR). 16-bit is CD-quality (~96 dB dynamic range); 24-bit (~144 dB) is the studio standard; 32-bit float covers far more than is audible and is used inside DAW mixers.
bone conduction Knochenleitung
Sound reaching the cochlea through the bones of the skull rather than through the outer and middle ear. It is why a recorded voice sounds wrong to its owner: in speaking, the speaker hears the airborne sound everyone else hears plus a bone-conducted component from the vocal folds, which favours low frequencies and makes the voice fuller from inside than from outside. The same path is used deliberately by bone-conduction headphones, which leave the ear canal open, and diagnostically in audiometry, where comparing bone- and air-conducted thresholds separates a conductive loss in the middle ear from a sensorineural one in the cochlea.Introduced in Chapter 1 — Communication, Frequency & Pitch.
bounce Bounce (Zusammenmischen, Mixdown)
To mix down a group of tracks (or apply a chain of processing) and write the result to a new track or file. The term survives from tape practice, where bouncing freed up tracks for further overdubs at the cost of an extra generation of noise.Introduced in Chapter 1 — Communication, Frequency & Pitch.
buffer size Puffergröße
The number of samples a real-time audio system processes per cycle. Smaller buffers reduce latency but raise the chance of dropouts (xruns) when the host CPU can’t keep up; larger buffers are safer but make monitoring while recording feel sluggish. Typical values: 32–128 samples for tracking, 256–1024 for mixing.Introduced in Chapter 9 — 3D Sound & Auralization.
bus line (also buss line) Bus (Summenbus / Sammelschiene)
A summing bus in a mixer: a signal path that gathers contributions from many input channels into a common output. A console may carry stereo busses, group busses, auxiliary effect busses, and a main mix bus, each addressable from every channel strip.Introduced in Chapter 5 — Miking & Recording.

C

cardioid Niere (Nierencharakteristik)
A heart-shaped microphone directivity pattern that is most sensitive on axis (0°), about 6 dB down at the sides (±90°) and has a deep null directly behind (180°). It is obtained by combining omni-directional and bi-directional elements in equal proportion.Introduced in Chapter 5 — Miking & Recording.
carrier Trägersignal (Träger)
In amplitude or frequency modulation, the high-frequency signal whose amplitude or frequency is varied by the modulator. Sidebands appear around the carrier at frequencies equal to the carrier ± integer multiples of the modulator.Introduced in Chapter 3 — Spectra, Modulation & Phase.
casting Besetzung (Casting)
The selection and assignment of sounds to the roles of narration, music and sound effect within an audio production — the audio counterpart of casting actors in film.Introduced in Chapter 4 — Sound Authoring & Casting.
cent Cent
A musical interval equal to 1/100 of an equally-tempered semitone, so 1200 cents fill an octave. Trained listeners may notice pitch differences of a few cents under favourable conditions, but ordinary discrimination varies strongly with frequency, level, duration, and context. The cent ratio is 21/1200 ≈ 1.0005777.Introduced in Chapter 1 — Communication, Frequency & Pitch.
center frequency Mittenfrequenz
The frequency at the geometric center of a band-pass or band-stop filter’s response, where attenuation (or boost) is at its extreme value. Together with bandwidth (or Q) it fully specifies a simple resonant section.Introduced in Chapter 7 — Effects & DSP.
cepstrum Cepstrum
The spectrum of the logarithm of a spectrum — the coinage reverses the first syllable of “spectrum,” and its horizontal axis is correspondingly called quefrency, measured in seconds rather than hertz. Taking the logarithm turns the product of an excitation and a resonance into a sum, which the second transform then pulls apart: a voice’s evenly spaced harmonic comb collapses to one sharp peak at the quefrency of its pitch period, well clear of the slowly varying spectral envelope that carries timbre and formant information. That separation makes the cepstrum a classical pitch estimator and the basis of source-filter analysis; its mel-scaled variant served for decades as the standard front end of speech recognition, before learned representations displaced it.Introduced in Chapter 3 — Spectra, Modulation & Phase.
clipping Clipping (Übersteuerung)
Distortion that occurs when a signal exceeds the maximum level a circuit or converter can represent, so that the waveform tops and bottoms are flattened. Clipping introduces high-order harmonic distortion and intermodulation components and is generally avoided in recording.Introduced in Chapter 2 — Intensity & Loudness.
clipping distortion
See clipping.Introduced in Chapter 6 — Digitization & Editing.
clave Clave
In Afro-Cuban music, a short asymmetric pattern of five onsets repeating over a two-bar cycle, played on a pair of hardwood sticks of the same name and acting as the timekeeper the whole ensemble orients to. The son and rumba claves differ by a single onset. More broadly, a timeline: the same role is filled by bell patterns across West Africa, and by the tresillo and cinquillo of the Caribbean.Playable, measurable, and comparable against its relatives on the Rhythm Playground bonus page.
col legno col legno (mit Holz)
A string-playing technique in which the wooden stick of the bow, rather than the hair, is struck or drawn across the string. The resulting sound is percussive, inharmonic and noise-like, used for special color by orchestral composers since the 17th century.Introduced in Chapter 3 — Spectra, Modulation & Phase.
comb filtering Kammfilterung
The notched-and-peaked frequency response produced when a signal is summed with a delayed copy of itself. Each spectral notch sits at a frequency where the delay equals an odd half-period, so identical sounds combine destructively. Heard when two microphones on one source mix together, when a speaker reflects off a nearby wall, and as the coloring effect of short (sub-20 ms) delays.
compact / desktop / bookshelf loudspeakers Kompaktlautsprecher / Desktop-Lautsprecher / Regal-Lautsprecher
Three overlapping but distinct terms for small loudspeakers. Compact is a size / marketing description: any small enclosure — desktop, bookshelf, satellite, small monitor, portable, or small surround speaker — can be marketed as compact. Desktop describes use and placement: a small speaker meant to sit beside a computer for near-field listening, often powered (active) with built-in amplifier, USB / Bluetooth inputs, and a front-panel volume knob. Bookshelf also describes use and placement — a small-to-medium hi-fi speaker meant to sit on stands, shelves, or furniture — and is usually passive, driven by a separate amplifier or AV receiver, though powered bookshelf speakers exist too.Introduced in Chapter 8 — Playback & Sound Check.
compander Kompander
A signal processor that combines a compressor on the way into a noisy channel with a complementary expander on the way out, extending the channel’s usable dynamic range. Used in tape noise reduction (Dolby, dbx) and in telephony codecs (μ-law, A-law).Introduced in Chapter 7 — Effects & DSP.
companding Kompandierung
The act of compressing the dynamic range of a signal before transmission or storage and expanding it again on playback, so that low-level detail is lifted clear of channel noise. Implemented by a compander.Introduced in Chapter 7 — Effects & DSP.
composer Komponist / Komponistin
The person who plans and arranges the sonic content of a work. Sonic uses the term broadly, covering both traditional music composition and the design of dialogue, ambience and effects for video, games, and interactive media.Introduced in Chapter 1 — Communication, Frequency & Pitch.
composition Komposition
The act, or the resulting work, of shaping sounds into a coherent whole — choosing what is heard, when and in what relation to other sensory material. In modern media production the composition includes silence, ambience and effects as well as music.Introduced in Chapter 1 — Communication, Frequency & Pitch.
compression Verdichtung (Überdruckphase)
In acoustics, the half of a sound wave in which air molecules are momentarily packed more densely than in the surrounding medium; the complement of rarefaction. (For dynamic-range compression see compressor.)Introduced in Chapter 1 — Communication, Frequency & Pitch.
compressor Kompressor
A dynamics processor that reduces the gain of a signal whenever its level exceeds a threshold, controlled by ratio, attack, and release parameters. Used to even out level differences, tighten transients and increase apparent loudness.Introduced in Chapter 6 — Digitization & Editing.
concert A Kammerton A (Konzert-A)
The reference pitch A4 ≈ 440 Hz, used to tune Western instruments to a common standard. 440 Hz was adopted as a recommendation by ISO in the 20th century and remains the most common modern tuning, though earlier centuries used a wide range of A4 frequencies and some orchestras still tune slightly sharper or flatter.Introduced in Chapter 1 — Communication, Frequency & Pitch.
condenser Kondensatormikrofon
A microphone whose transducer is a capacitor formed by a thin charged diaphragm and a fixed back plate; sound pressure varies the spacing and thus the capacitance, generating an output voltage. Condenser elements are prized for their flat response and fast transient handling, but the capsule has to carry a charge. An externally polarized condenser draws that charge from a supply, commonly derived from phantom power; an electret condenser holds a permanent charge in its diaphragm or back plate and needs none, though its internal impedance converter still has to be powered.Introduced in Chapter 5 — Miking & Recording.
constructive konstruktive Interferenz
Constructive interference: the addition of two signals whose phases agree, producing an output larger than either component. Two in-phase sine waves of equal amplitude sum to a sine of twice the amplitude.Introduced in Chapter 3 — Spectra, Modulation & Phase.
convolution Faltung
The operation that combines two signals into a third by flipping one of them, sliding it by a shift t, multiplying the two point-by-point, and summing the products: y(t) = Σ x(τ)·h(t−τ). It is the defining operation of a linear time-invariant system: the output of any filter, loudspeaker, or room is its input convolved with that system’s impulse response. Convolution is commutative — x ⊛ h = h ⊛ x — and corresponds to multiplication in the frequency domain, which is why long convolutions are computed through a fast Fourier transform and its inverse rather than sample by sample.Explored on the Convolution Explorer; applied as convolution reverb in Chapter 7 — Effects & DSP and Chapter 9 — 3D Sound & Auralization.
convolution reverb Faltungshall
Reverberation produced by convolving a dry signal with the measured or synthesized impulse response of a real or modeled space, rather than by a network of delays and filters designed to imitate one. Because the impulse response captures a particular room heard from a particular position, convolution reverb reproduces that space’s early-reflection pattern and decay directly; the cost is that the room cannot then be reshaped without a new impulse response.Introduced in Chapter 9 — 3D Sound & Auralization; the underlying operation is explored on the Convolution Explorer.
critical band Frequenzgruppe (kritische Bandbreite)
The frequency range within which the ear treats two components as competing for the same place, rather than resolving them separately. Physically it corresponds to a stretch of the basilar membrane: tones close enough to excite overlapping regions are not heard as two, but as one fluctuating sound — slowly, as beats, or fast enough to be heard as roughness. Tones further apart fall in different critical bands and separate into two pitches. The bandwidth is roughly a fifth to a sixth of its own centre frequency across most of the audio range, so it is narrow low down and wide up high, which is why a semitone sounds muddy in the bass and clear in the treble. The same width governs masking, and so the bit allocation of perceptual codecs, which spend nothing on what a louder neighbour in the same band has already hidden.Introduced in Chapter 1 — Communication, Frequency & Pitch.
cross-talk Übersprechen
Unwanted leakage of a signal from one channel into another, whether through electrical coupling, mechanical coupling, or — in loudspeaker stereo — the acoustic path from each speaker to the listener’s opposite ear. Acoustic cross-talk is what prevents binaural recordings from imaging correctly over speakers.Introduced in Chapter 9 — 3D Sound & Auralization.
Cross-talk cancellation Übersprechkompensation (Crosstalk-Cancellation)
A technique for delivering binaural signals through loudspeakers by feeding each speaker a phase-and-time-corrected anti-signal that cancels its arrival at the opposite ear. Effective only within a small sweet spot and for a known head position.Introduced in Chapter 9 — 3D Sound & Auralization.
cut-off
See cut-off frequency.Introduced in Chapter 7 — Effects & DSP.
cut-off frequency Grenzfrequenz (Eckfrequenz)
The frequency that marks the boundary between a filter’s passband and stopband, conventionally defined as the point of −3 dB attenuation (half-power). For a low-pass filter frequencies above this are progressively attenuated; for a high-pass filter, those below.Introduced in Chapter 7 — Effects & DSP.

D

D-A
See digital-analog converter.Introduced in Chapter 4 — Sound Authoring & Casting.
dB RMS dB RMS (Effektivwertpegel)
A decibel measurement of a signal’s root-mean-square (average power) level rather than its instantaneous peak. RMS tracks perceived loudness more closely than a peak reading does, and it is what VU-style averaging meters approximate. Modern programme-loudness standards go further: the ITU-R BS.1770 family behind LUFS applies a frequency weighting and gates out quiet passages before integrating, so an RMS figure and a loudness figure for the same material need not agree.Introduced in Chapter 6 — Digitization & Editing.
dB sound pressure level Schalldruckpegel
A logarithmic measure of acoustic sound pressure referenced to 20 µPa (the nominal threshold of human hearing at 1 kHz): dB SPL = 20·log10(P/20 µPa). 0 dB SPL is the reference; conversational speech is about 60 dB SPL and a jet at 30 m about 120 dB SPL.Introduced in Chapter 2 — Intensity & Loudness.
dB SPL
See dB sound pressure level.Introduced in Chapter 2 — Intensity & Loudness.
dB SPL meter
See sound pressure level meter.Introduced in Chapter 2 — Intensity & Loudness.
dBFS dBFS (Dezibel bezogen auf Vollaussteuerung)
Decibels relative to full scale — the reference level used to measure digital audio. 0 dBFS is the maximum representable sample value; all undistorted digital signals sit at or below 0 dBFS, so meter readings in dBFS are always negative or zero. Distinct from dB SPL (acoustic) and dBu / dBV (analog electrical).
destructive interference destruktive Interferenz
The addition of two signals whose phases oppose, producing an output smaller than either — or, when they are equal in amplitude and 180° out of phase, complete cancellation. Destructive interference produces the comb-filter notches that follow short delays and room reflections.Introduced in Chapter 3 — Spectra, Modulation & Phase.
diaphragm Membran
The thin, flexible membrane of a microphone (or loudspeaker) that vibrates in response to sound pressure (or electrical drive). Its mass, tension and area set the element’s sensitivity, frequency response and transient behavior.Introduced in Chapter 5 — Miking & Recording.
diffusion Diffusion (Schallstreuung)
The scattering of reflected sound, by which distinct reflections lose their individual identity and direction. An irregular or broken surface returns one arriving wavefront as many weaker ones spreading in all directions, rather than as a single mirror-like bounce. Diffusion is what turns early reflections into late reverberation: each reflection spawns further reflections, so their number grows roughly as the square of elapsed time until arrivals overlap too densely to be separated either by the ear or by a measurement. That rate is the echo density, and the moment it passes the point of resolvability is the mixing time — after it, a room is described statistically, by how fast it decays, rather than reflection by reflection. A diffusive room reaches that state sooner and sounds smoother; a room of flat parallel surfaces reaches it late or never, and keeps audible flutter. The word carries a second, related sense in reverb design, where a diffusion control sets how quickly an algorithm builds echo density.Introduced in Chapter 7 — Utilities & Effects.
digital audio workstation (DAW) Digitale Audio-Workstation (DAW)
An integrated software environment for recording, editing, mixing, and mastering digital audio, typically combining multitrack arrangement, plug-in hosting, MIDI sequencing and surround/ambisonic routing. Familiar examples include Pro Tools, Logic Pro, Reaper, and Ableton Live.Introduced in Chapter 4 — Sound Authoring & Casting.
digital foldover
See aliasing.Introduced in Chapter 6 — Digitization & Editing.
digital signal processing (DSP) digitale Signalverarbeitung (DSP)
Any computation performed on a discrete-time signal (a stream of numerical samples) to modify, analyze, or synthesize sound. Typical DSP operations on audio include filtering (low-pass, high-pass, parametric EQ), mixing, gain and dynamics processing (compressors, limiters), modulation effects (chorus, flanger, tremolo), reverberation and convolution, time-stretching and pitch-shifting, spectral analysis (FFT), and synthesis (oscillators, sampling, granular). DSP can run offline on stored samples or in real-time on streaming audio via dedicated DSP chips, CPU SIMD code, or GPU shaders. Introduced in Chapter 1 — Communication, Frequency & Pitch; revisited as the explicit block between storage and D-A in Chapter 5 — Miking & Recording and as the primary subject of Chapter 7 — Effects & DSP.
digital-analog converter Digital-Analog-Wandler (D/A-Wandler)
A circuit that converts a stream of digital sample codes back into a continuous analog voltage, typically followed by a reconstruction low-pass filter to remove image spectra above the Nyquist limit. Abbreviated D-A or DAC; the inverse of an A-D converter.Introduced in Chapter 1 — Communication, Frequency & Pitch.
digitization Digitalisierung
The conversion of a continuous analog signal into a discrete-time, discrete-amplitude digital representation by sampling and quantization. The combined operations of an A-D converter.Introduced in Chapter 4 — Sound Authoring & Casting.
directivity index Bündelungsmaß (Richtwirkungsindex)
Ratio (in decibels) of a directional microphone's on-axis sensitivity to the average sensitivity over a full 4π sphere; equivalently, how much more sensitive the mic is to sound arriving from its target direction than to a diffuse field. 0 dB = omni-directional; ≈ 4.77 dB = cardioid; ≈ 5.71 dB = super-cardioid; ≈ 6.02 dB = hyper-cardioid. Higher DI = tighter pickup pattern, more rejection of off-axis sound. Introduced in Chapter 5 — Miking & Recording.
directivity pattern Richtcharakteristik
The polar plot showing a microphone’s (or loudspeaker’s) sensitivity as a function of the angle of incidence of sound. Common patterns include omni-directional, cardioid, super- and hyper-cardioid, and bi-directional (figure-8).Introduced in Chapter 5 — Miking & Recording.
dispersion Dispersion
The dependence of a wave’s propagation speed on its frequency. In a dispersive medium the phase velocity and the group velocity differ, so a wave packet changes shape as it travels and its crests move through its own envelope. Sound in air is very nearly non-dispersive across the audio band — which is why an orchestra heard from the back of a hall still arrives in time — but stiff strings and bars, shallow water, and acoustic waveguides all disperse. The signal-processing counterpart is a frequency-dependent group delay, as produced by an all-pass filter.Demonstrated on the Phase & Group Velocity bonus page.
distant miking Distanzmikrofonierung (Hauptmikrofon-Aufnahme)
A recording technique in which the microphones are placed well back from the source, capturing the blended sound of an ensemble together with the room’s reverberation. Contrasted with spot miking.Introduced in Chapter 5 — Miking & Recording.
Doppler effect Dopplereffekt
The shift in observed frequency caused by relative motion between a source and a listener: the pitch rises as the source approaches and falls as it recedes, most familiar as the drop in a passing siren. The shift factor is C / (C + vradial), with C the speed of sound; because it scales frequency itself, a modulated source has its entire spectrum shifted — both carrier and deviation — not just its center.Demonstrated on the Sirens & FM bonus page; related to Chapter 9 — 3D Sound & Auralization.
down-sampling Abwärtsabtastung (Downsampling)
The process of reducing a signal’s sample rate, preceded by a low-pass filter set below the new Nyquist frequency to prevent aliasing. Used to deliver mastered files at a target rate or to reduce computational load.Introduced in Chapter 7 — Effects & DSP.
dry / wet signal Dry/Wet (Originalsignal / Effektsignal)
In an effects chain, dry is the unprocessed source and wet is the processed copy that comes out of the effect (e.g. the reverb tail, the echoed repeats). The wet/dry mix control balances how much of each is heard in the output — 100 % dry bypasses the effect, 100 % wet hides the source under the effect alone.Introduced in Chapter 7 — Effects & DSP.
dummy head recording Kunstkopfaufnahme
A binaural recording made with microphones mounted at the ear-canal entrances of an anatomical head-and-torso replica (e.g. Neumann KU 100). The dummy head imposes a realistic HRTF on incoming sound, yielding strong externalization on headphone playback.Introduced in Chapter 9 — 3D Sound & Auralization.
dynamic dynamisches Mikrofon (Tauchspulmikrofon)
A microphone whose transducer is a coil of wire (the voice coil) attached to the diaphragm and suspended in a magnetic field; diaphragm motion induces a voltage across the coil. Dynamic mics are rugged, need no polarizing voltage, and tolerate very high sound-pressure levels.Introduced in Chapter 5 — Miking & Recording.
dynamic range Dynamikumfang (Dynamikbereich)
The span from the quietest sound that can be detected to the loudest that can be tolerated, expressed in dB. Human hearing reaches roughly 120 dB at the ear’s most sensitive frequencies; audio equipment is characterized by the analogous span between its noise floor and its clipping ceiling.Introduced in Chapter 2 — Intensity & Loudness.

E

earcon Earcon
An interface sound that signifies by convention rather than resemblance: a short abstract motif, such as two rising notes for success, which means nothing until learned and thereafter means it exactly. The name puns on icon and ear. Earcons can be built in families — a shared rhythm marking a category while pitch or timbre distinguishes its members — so unlike auditory icons they scale to as many messages as needed, at the cost of having to be taught.Introduced in Chapter 4 — Sound Authoring & Casting.
early reflections frühe Reflexionen (Erstreflexionen)
The first echoes that reach a listener after the direct sound, arriving within roughly the first 50–100 ms via single bounces off walls, ceiling and nearby surfaces. They convey strong cues about room size, source distance and listener position.Introduced in Chapter 8 — Playback & Sound Check.
echo Echo
A reflected sound that arrives at a listener with enough delay (roughly more than 50 ms for most sources, less for transients) that the ear perceives it as a distinct repetition of the direct sound rather than fusing with it as part of the reverberant tail.Introduced in Chapter 7 — Effects & DSP.
Euclidean rhythm euklidischer Rhythmus
The pattern E(k,n) that spreads k onsets as evenly as possible among n equal steps of a cycle. The construction is Bjorklund’s algorithm, written for the timing system of a particle accelerator, and it is a disguised form of Euclid’s procedure for the greatest common divisor. Toussaint observed that a striking number of traditional rhythms are exactly Euclidean — E(3,8) is the tresillo, E(5,8) the cinquillo — while others, the son clave among them, sit conspicuously one step away.Generated and tested on the Rhythm Playground bonus page.
edit
See editing.Introduced in Chapter 1 — Communication, Frequency & Pitch.
editing (Ton-)Schnitt / Audio-Editing
The craft of assembling, trimming and refining recorded audio. Modern editing is performed graphically on waveform displays inside a DAW, allowing sample-accurate cuts, crossfades, and time-warping.Introduced in Chapter 4 — Sound Authoring & Casting.
effect send Effekt-Send (Aux-Weg)
A mixer output, tapped from each channel by an auxiliary control, that routes a portion of the channel’s signal to an effects processor (reverb, delay, etc.). The processed return is then mixed back into the main bus.Introduced in Chapter 5 — Miking & Recording.
electrostatic
See condenser.Introduced in Chapter 5 — Miking & Recording.
environmental context akustische Umgebung (Umgebungskontext)
The acoustic surround in which a sound is heard — primarily its pattern of reflections and reverberation. Environmental context is what allows a listener to recognize a recording as taking place in a cathedral, a tiled bathroom or an open field.Introduced in Chapter 9 — 3D Sound & Auralization.
equal-loudness contours Kurven gleicher Lautstärke (Isophone)
A family of curves (Fletcher–Munson, later Robinson–Dadson, now ISO 226) that show what SPL a sine tone needs at each frequency to be heard as equally loud as a 1 kHz reference. The contours bow upward at the extremes, especially the low end — which is why hi-fi systems used to have a loudness button that boosted bass at low listening levels.Introduced in Chapter 2 — Intensity & Loudness.
excitation Anregung
The transfer of energy that sets a vibrating system into motion — bowing or plucking a string, blowing across a reed, striking a drumhead, or driving a loudspeaker cone. The excitation determines the attack of the resulting sound and which modes of the resonator are activated.Introduced in Chapter 3 — Spectra, Modulation & Phase.

F

filter (plural filters) Filter
Linear processes that selectively boost or attenuate spectral regions of a signal. Common audio filters include low-pass, high-pass, band-pass, band-stop, shelving, and parametric peaking sections, used for tone shaping, anti-aliasing, and noise removal.Introduced in Chapter 3 — Spectra, Modulation & Phase.
floating-point numbers Gleitkommazahlen
A numeric representation that stores each value as a mantissa and an exponent, giving a very wide dynamic range at the cost of slightly non-uniform precision. Modern DAWs run their internal signal path in 32- or 64-bit float, so intermediate sums can exceed 0 dBFS without clipping.Introduced in Chapter 2 — Intensity & Loudness.
Foley sound Foley-Geräusche (Geräuschemacherei)
Everyday sound effects — footsteps, cloth rustles, prop handling — that are performed and recorded in sync with picture to replace or augment the production track. Named after sound editor Jack Foley (1891–1967).Introduced in Chapter 4 — Sound Authoring & Casting.
frequency Frequenz
The number of complete cycles a periodic waveform completes per second, measured in hertz (Hz). For a pure tone, frequency is the principal physical correlate of perceived pitch.Introduced in Chapter 1 — Communication, Frequency & Pitch.
frequency modulation (FM) Frequenzmodulation (FM)
A technique in which the instantaneous frequency of a carrier is varied in proportion to a second, modulating signal. At sub-audio modulation rates the ear follows the excursion and hears vibrato or a siren’s glide; once the modulator reaches audio rates the sweep is too fast to track and is heard instead as a new timbre, because the modulation spreads the carrier’s energy into a comb of sidebands spaced at the modulator frequency, with the number of significant sidebands growing as the modulation index rises. It is the basis of FM synthesis and of the wail, yelp, and phaser siren voices.Introduced in Chapter 3 — Spectra, Modulation & Phase; explored on the Sirens & FM bonus page.
frequency response Frequenzgang
The way the gain (and phase) of a system varies with frequency, usually displayed as a magnitude curve on a logarithmic frequency axis. The frequency response of microphones, loudspeakers and rooms shapes the timbre of everything that passes through them.Introduced in Chapter 1 — Communication, Frequency & Pitch.
fundamental Grundton
The lowest-frequency partial of a harmonic complex tone. Its frequency normally determines the perceived pitch, and its period is the period of the overall waveform.Introduced in Chapter 3 — Spectra, Modulation & Phase.
fundamental tracking Fundamentalverfolgung (Effekt des fehlenden Grundtons / Residualton)
The perceptual phenomenon in which the ear assigns a pitch corresponding to the fundamental frequency of a harmonic series even when that fundamental is missing from the actual spectrum — also called residue pitch or the missing fundamental. The brain reconstructs the implied fundamental from the spacing of the higher harmonics, which is why a small loudspeaker can convey the pitch of a bass note it cannot physically reproduce.Introduced in Chapter 3 — Spectra, Modulation & Phase.

G

gain staging Gain-Staging (Pegelstufenabgleich)
The practice of setting the gain at each stage of an audio chain so that every stage sits at a sensible level — well above the noise floor, well below the clipping ceiling. Good gain staging keeps a consistent nominal level (e.g. −18 dBFS RMS) through preamps, converters, plugins, and busses, and is the single biggest determinant of recording quality after microphone choice.Introduced in Chapter 5 — Miking & Recording.
graphic equalizer grafischer Equalizer
A multi-band equalizer whose fixed-frequency bands are controlled by linear sliders, so that the slider positions form a graphic picture of the frequency response. Typically built from a bank of band-pass sections covering octaves or third-octaves.Introduced in Chapter 3 — Spectra, Modulation & Phase.
graphic equalizers (EQs)
See graphic equalizer.Introduced in Chapter 7 — Effects & DSP.
group velocity Gruppengeschwindigkeit
The speed at which the envelope of a wave packet travels — and with it the packet’s energy and whatever information it carries: vg = dω/dk, the slope of the dispersion relation. In a non-dispersive medium it equals the phase velocity and the packet travels rigidly; under dispersion the two separate, and individual crests can be seen to appear at one edge of the packet, cross it, and vanish at the other. Deep-water waves are the textbook case, with vg = ½ vp. The filter counterpart is group delay, −dφ/dω.Demonstrated on the Phase & Group Velocity bonus page.

H

half-step Halbton
The smallest interval in twelve-tone equal temperament: a frequency ratio of 21/12 ≈ 1.0595, or 100 cents. Also called a semitone; it is the distance from one piano key to the next, white or black.Introduced in Chapter 3 — Spectra, Modulation & Phase.
harmonics Obertöne
Partials whose frequencies are integer multiples of the fundamental frequency. Harmonic spectra are characteristic of strings, brass and the voice, and are the basis of Western tonal pitch perception.Introduced in Chapter 3 — Spectra, Modulation & Phase.
The direction-dependent linear filter, measured separately for each ear, that describes how a free-field source is shaped by the listener’s torso, head and pinnae on the way to the eardrum. Convolving an anechoic signal with a pair of HRTFs is the foundation of binaural 3D audio rendering.Introduced in Chapter 9 — 3D Sound & Auralization.
A block of metadata at the beginning of a sound file describing how to interpret the audio data that follow: sample rate, channel count, bit depth, encoding and (in WAV/BWF, AIFF, FLAC) optional chunks for cue points, markers and broadcast metadata.Introduced in Chapter 6 — Digitization & Editing.
headroom Aussteuerungsreserve (Headroom)
The margin, in decibels, between the nominal operating level of a system and the level at which it begins to clip. Adequate headroom accommodates transient peaks above the average signal level without distortion.Introduced in Chapter 6 — Digitization & Editing.
hearing Gehör
The perceptual process by which acoustic pressure variations at the eardrum are transduced into neural activity and interpreted by the brain as sound. It encompasses peripheral mechanics (outer, middle and inner ear) and central auditory processing.Introduced in Chapter 1 — Communication, Frequency & Pitch.
hertz (Hz) Hertz
The SI unit of frequency, equal to one cycle per second. Abbreviated Hz; named after Heinrich Hertz (1857–1894). The audible range for healthy young humans is conventionally 20 Hz to 20 kHz. Note: the unit symbol Hz is capitalised because it is named after Heinrich Hertz; the spelled-out form hertz is lowercase.Introduced in Chapter 1 — Communication, Frequency & Pitch.
high-pass filters (HPF) Hochpass(filter)
A filter that passes frequencies above its cut-off and attenuates those below; useful for removing rumble, handling noise, and DC offsets. Specified by cut-off frequency and slope (often 6, 12 or 24 dB/octave).Introduced in Chapter 7 — Effects & DSP.

I

immersion Immersion
The subjective sense of being enveloped by, and present within, a virtual sonic environment, such that the listener’s attention is drawn away from the surrounding real world. Immersion is enhanced by accurate spatial cues, low-latency interactivity and a wide bandwidth playback chain.Introduced in Chapter 1 — Communication, Frequency & Pitch.
impedance Impedanz (Scheinwiderstand)
The frequency-dependent opposition a circuit presents to alternating current, measured in ohms. Microphone, line and loudspeaker stages are designed for specific source-to-load impedance ratios; mismatches cause level loss, frequency-response errors or noise.Introduced in Chapter 5 — Miking & Recording.
impulse response Impulsantwort
What a linear time-invariant system — a filter, a loudspeaker, a room — puts out when a single instantaneous click or tap goes in. Because any signal can be regarded as a dense train of scaled, shifted impulses, the impulse response completely characterizes the system: the output for any input is that input convolved with it. A room’s impulse response shows the direct sound as a first spike, the discrete early reflections as later ones, and the diffuse decaying tail as late reverberation; its reverberation time can be read straight off the decay. Its frequency-domain equivalent is the frequency response.Applied in Chapter 7 — Effects & DSP and Chapter 9 — 3D Sound & Auralization; its role in convolution is shown on the Convolution Explorer.
inharmonic inharmonisch
Describing partials whose frequencies are not integer multiples of a common fundamental. Inharmonic spectra are typical of bells, gongs, drums, and stiff piano strings, and they make pitch perception less definite.Introduced in Chapter 3 — Spectra, Modulation & Phase.
intensity Intensität
The acoustic power crossing a unit area perpendicular to the direction of propagation, measured in watts per square meter. Intensity falls as the inverse square of distance from a point source in a free field.Introduced in Chapter 2 — Intensity & Loudness.
interaural level difference (ILD) interaurale Pegeldifferenz
Also written interaural intensity difference. The difference in sound level between the two ears, produced by the head’s acoustic shadow at frequencies above roughly 1.5 kHz. Together with the interaural time difference it is a primary cue for the lateral location of a sound source.Introduced in Chapter 9 — 3D Sound & Auralization.
interaural time differences interaurale Laufzeitdifferenz
The difference in time of arrival of a sound at the two ears (ITD), at most about 660 µs for a source to one side. ITD dominates lateral localization at low frequencies, where head-shadow level differences are small.Introduced in Chapter 9 — 3D Sound & Auralization.
inverse FFT inverse FFT (Rücktransformation)
The mathematical inverse of the fast Fourier transform: it reconstructs a time-domain signal from a set of complex frequency-domain samples. Block-based DSP such as fast convolution, phase vocoding and spectral effects depend on a forward FFT followed by an inverse FFT after spectral modification.Introduced in Chapter 7 — Effects & DSP.

J

just noticeable difference (JND) Unterschiedsschwelle (eben merklicher Unterschied)
The smallest change in a stimulus that a listener can reliably detect, expressed as Δx or as a fraction of x. For frequency at moderate levels the JND is roughly 0.2–0.5 %; for level it is about 1 dB.Introduced in Chapter 1 — Communication, Frequency & Pitch.

K

kHz
See kilohertz.Introduced in Chapter 1 — Communication, Frequency & Pitch.
kilohertz Kilohertz
A unit of frequency equal to 1000 hertz, abbreviated kHz. Convenient for the upper part of the audible range: 1 kHz is roughly two octaves above middle C, and the human hearing range extends to about 20 kHz.Introduced in Chapter 1 — Communication, Frequency & Pitch.

L

lapel microphones Lavaliermikrofon (Ansteckmikrofon)
Small omni- or cardioid microphones designed to be clipped to a presenter’s clothing, near the sternum, for hands-free speech pickup; also called lavalier mics. Used extensively in broadcast, theatre and conference work, usually paired with a body-pack wireless transmitter.Introduced in Chapter 5 — Miking & Recording.
late reverberation Spätnachhall (später, dichter Nachhall)
The diffuse, exponentially-decaying tail of a room’s impulse response that follows the discrete early reflections, formed when individual reflections become too numerous and closely-spaced to be heard separately. Sometimes called dense reverberation in reference to this temporal density.Introduced in Chapter 7 — Effects & DSP.
latency Latenz
The delay between an audio event entering a system and emerging from it. In DAWs it is dominated by the converter buffer size and any plugins that look ahead; in network audio it adds transmission and jitter-buffer delay. Latencies above ~10 ms are noticeable when monitoring one’s own playing; above ~30 ms they become disorienting.
lateralized Im-Kopf-Lokalisation (lateralisiert)
Heard as inside or at the edge of the head rather than out in space. Stereo signals presented over headphones without HRTF processing tend to be lateralized between the ears, not externalized into the surrounding scene.Introduced in Chapter 9 — 3D Sound & Auralization.
line level Line-Pegel
The nominal signal level used to interconnect mixers, recorders, processors, and amplifiers. Professional line level is +4 dBu (about 1.23 V RMS); consumer line level is −10 dBV (about 0.316 V RMS).Introduced in Chapter 5 — Miking & Recording.
loops Loops (Schleifen)
Short audio segments that play back end-to-end repeatedly to extend a sound — a drum groove, a synth pad, an ambience bed. Seamless looping requires matched start and end points, ideally at zero crossings and on a musical phrase boundary.Introduced in Chapter 4 — Sound Authoring & Casting.
lossy compression verlustbehaftete Kompression
Data compression that discards perceptually less important information to achieve a smaller file, so that the decoded signal is not bit-identical to the original. MP3, AAC, Opus, and Ogg Vorbis are common lossy audio codecs.Introduced in Chapter 7 — Effects & DSP.
loudness Lautheit
The subjective perception of a sound’s intensity. Loudness depends on sound pressure level but also on frequency content, duration and context, as captured by the equal-loudness contours and by standards such as LUFS (ITU-R BS.1770).Introduced in Chapter 2 — Intensity & Loudness.
loudness units full-scale (LUFS) LUFS (Lautheit relativ zu digitalem Vollaussteuerungspegel)
The contemporary loudness measurement standardized by ITU-R BS.1770, expressed in units of dB relative to digital full scale but weighted to the ear’s frequency response and integrated over time. Streaming platforms target programme loudness around −14 LUFS (Spotify, Apple Music) or −23 LUFS (EBU broadcast). Distinct from dB SPL (acoustic) and from RMS (un-weighted).Introduced in Chapter 2 — Intensity & Loudness.
low-frequency oscillator (LFO) Niederfrequenzoszillator (LFO)
A slow oscillator — typically below 20 Hz — used as a modulation source rather than as a sound itself. Its output is fed to another parameter to create tremolo (amplitude modulation), vibrato (frequency modulation), auto-pan, filter sweeps, and other periodic time-varying effects.Introduced in Chapter 3 — Spectra, Modulation & Phase.
low-pass filter (LPF) Tiefpass(filter)
A filter that passes frequencies below its cut-off and attenuates those above; used for anti-aliasing, reconstruction, and tone shaping. Specified by cut-off frequency and slope.Introduced in Chapter 7 — Effects & DSP.
low-pass filtering
See low-pass filter.Introduced in Chapter 6 — Digitization & Editing.

M

masking (auditive) Verdeckung / Maskierung
A perceptual phenomenon in which a louder sound (the masker) makes a quieter sound near it in frequency or time inaudible. Lossy audio codecs (MP3, AAC, Vorbis) exploit masking to discard signal components the ear would not have heard, achieving compression with little perceived loss.Introduced in Chapter 0 — Welcome.
master gain Master-Pegel (Summen-Pegel)
The final output level control of a mixer or sub-mix, applied after all channel and bus processing. It scales the entire mix without altering the relative balance between sources.Introduced in Chapter 5 — Miking & Recording.
McGurk effect McGurk-Effekt
An audiovisual illusion in which watching a talker changes what is heard. Dubbing the recording of one spoken syllable onto video of a mouth articulating another commonly yields a third syllable that neither channel contains: the standard case pairs an audio /ba/ with a visual /ga/ and is heard as “da.” Closing the eyes restores the recorded syllable at once, and the fusion returns on opening them — knowing the trick does not dissolve it. The effect is evidence that a percept is not a read-out of one sense but an inference fused across several, and it sets a practical bound on how far a soundtrack can diverge from the picture before a viewer hears the seam. Reported by Harry McGurk and John MacDonald in 1976.Introduced in Chapter 1 — Communication, Frequency & Pitch.
medium Medium (Übertragungsmedium)
A general term for whatever carries a signal between source and listener. In acoustics it is the physical material (air, water, a wall) that propagates the pressure wave; in production it is the storage substrate (tape, disk, file) holding the recording; in communication it is the channel (radio, telephony, streaming) by which the signal reaches its audience.Introduced in Chapter 1 — Communication, Frequency & Pitch.
MIDI MIDI
Musical Instrument Digital Interface — a control protocol (not audio) for exchanging note-on, note-off, pitch-bend, control-change, and clock messages between electronic instruments, computers, and DAWs. MIDI carries performance data, not sound; a MIDI file describes what was played, and a synthesizer or sample player turns it into audio.Introduced in Chapter 7 — Effects & DSP.
miniature microphone Miniaturmikrofon
A very small microphone capsule, usually omni-directional, designed for concealed placement on talent, instruments or props; widely used in theatre, film and broadcast.Introduced in Chapter 5 — Miking & Recording.
mixer Mischpult
A device, hardware or software, that sums multiple audio signals into one or more output busses while providing per-channel level, panning, equalization, and routing. See also audio mixing console.Introduced in Chapter 1 — Communication, Frequency & Pitch.
modulation index Modulationsindex
In frequency modulation, the ratio I = d / fm of the peak frequency deviation to the modulator rate. It governs how many sidebands carry significant energy and how bright the resulting timbre is: the amplitude of the n-th sideband pair is the Bessel function |Jn(I)|, and near I ≈ 2.405 the carrier itself vanishes. Low index reads as vibrato; high index builds an instrument.Introduced in Chapter 3 — Spectra, Modulation & Phase; explored on the Sirens & FM bonus page.
monitoring Abhören / Monitoring
Listening to audio during recording, mixing or live performance via dedicated playback monitors (speakers or headphones) selected for accurate, neutral reproduction. Studio monitoring deliberately avoids the coloration of consumer hi-fi so that production decisions transfer well to other systems.Introduced in Chapter 6 — Digitization & Editing.
mono / stereo / surround Mono / Stereo / Surround
Mono carries one audio channel — all listeners hear the same signal. Stereo carries two channels (left/right) that together place sources between the speakers and convey ensemble width. Surround formats (5.1, 7.1, Dolby Atmos, ambisonics) extend the idea to more speakers around and above the listener, supporting immersive playback.
multimedia loudspeakers
See compact / desktop / bookshelf loudspeakers.Introduced in Chapter 8 — Playback & Sound Check.
music Musik
Organized sound intended to be experienced for its expressive, structural or aesthetic qualities. As a cast category in Sonic it encompasses both scored cues and atmospheric beds that support a media narrative.Introduced in Chapter 4 — Sound Authoring & Casting.

N

narration Erzählung (Sprecherkommentar / Voice-over)
Spoken voice, usually addressed directly to the listener, that conveys information or guides the listener through a media work. As a cast category, narration is treated separately from in-scene dialogue and from sung music.Introduced in Chapter 4 — Sound Authoring & Casting.
narrowcasting Narrowcasting
The selective delivery of media to and from chosen parties in a shared virtual space — a middle ground between broadcasting (to everyone) and unicasting (to a single recipient). In a scene of several sources and several sinks, each source can be muted (silenced for all) or soloed (left as the only one heard), and each sink can be deafened (blocked from hearing a given source) or attended (focused on a chosen source); these four operators let a participant sculpt whom they hear and by whom they are heard. Mute and solo act on the source side, deafen and attend on the sink side.Introduced in Chapter 9 — 3D Sound & Auralization; explored in the narrowcasting sandbox.
natural hearing natürliches Hören
Listening to acoustic events directly in everyday environments — ears uncovered, head free to move, multiple sensory cues available — as contrasted with the mediated experience of headphone or loudspeaker playback.Introduced in Chapter 1 — Communication, Frequency & Pitch.
noise floor Rauschboden (Rauschpegel / Grundrauschen)
The lowest signal level that can be distinguished from background noise in a system or environment. Sounds below the noise floor are masked; the floor sets the lower bound of the system’s usable dynamic range.Introduced in Chapter 2 — Intensity & Loudness.
noise-like rauschartig (geräuschhaft)
Containing a broadband, aperiodic component that lacks a clear pitch. Noise-like elements appear in breath, bow scrape, percussion attacks, and sibilants, and contribute strongly to a sound’s identity even when they are spectrally subordinate to tonal partials.Introduced in Chapter 3 — Spectra, Modulation & Phase.
nominal level Nennpegel
The reference operating level around which a signal chain is designed to run, with the noise floor well below and headroom above. In analog work it is usually marked 0 VU; in digital work a target such as −18 or −20 dBFS is chosen to leave room for peaks. Often confused with nominal gain, which would be the gain that puts a typical input at this level — level is the state of the signal, gain is the operation that brings it there.Introduced in Chapter 6 — Digitization & Editing.
nominal gain
See nominal level.Introduced in Chapter 6 — Digitization & Editing.
normalization Normalisierung (Pegelnormalisierung)
A gain-scaling operation that lifts a file so its peak (or its measured loudness) hits a chosen target — typically 0 dBFS or a LUFS target. Depending on the editor it can be applied destructively (the samples themselves are rewritten) or non-destructively (a clip-gain or mixer-fader instruction). Normalization changes level but not dynamics; the ratio of loud-to-quiet within the file is unchanged.Introduced in Chapter 7 — Effects & DSP.
Nyquist theorem Nyquist-Theorem (Abtasttheorem)
The Nyquist–Shannon sampling theorem: a band-limited signal can be perfectly reconstructed from its samples provided the sampling rate strictly exceeds twice the highest frequency present. Components at or above half the sample rate alias into the audible band and must be removed by an anti-alias filter.Introduced in Chapter 6 — Digitization & Editing.

O

octave Oktave
The musical interval corresponding to a 2:1 frequency ratio. Notes an octave apart share a pitch class and are heard as the same note in different registers.Introduced in Chapter 1 — Communication, Frequency & Pitch.
omni-directional Kugelcharakteristik (omnidirektional)
A microphone directivity pattern with essentially equal sensitivity to sound arriving from any direction. Omnis are inherently free of proximity effect and tend to have the smoothest low-frequency and off-axis responses.Introduced in Chapter 5 — Miking & Recording.
oscillation Schwingung
One complete back-and-forth cycle of a vibrating quantity — air pressure, a string’s displacement, or a loudspeaker cone’s position. The number of oscillations per second is the frequency, measured in hertz.Introduced in Chapter 1 — Communication, Frequency & Pitch.
oscillator Oszillator
A circuit or algorithm that generates a periodic waveform — sine, square, triangle, sawtooth or arbitrary — at a controllable frequency. Oscillators are the source elements of synthesizers, the reference for test equipment, and the carriers for modulation.Introduced in Chapter 3 — Spectra, Modulation & Phase.
ostinato Ostinato
A short musical figure that is repeated persistently throughout a passage, often forming the rhythmic or harmonic foundation of a piece. In media scoring, an ostinato bed provides continuity beneath dialogue or shifting visuals.Introduced in Chapter 4 — Sound Authoring & Casting.
outboard Outboard-Geräte (externe Hardware)
External hardware processors — preamps, equalizers, compressors, reverbs — that are patched into a console’s signal path via inserts or auxiliary sends, as distinct from in-the-box software plug-ins.Introduced in Chapter 7 — Effects & DSP.

P

polymeter Polymetrik
Two or more parts sharing one pulse but grouping it into cycles of different lengths — a seven-step pattern against a sixteen-step one, both counting the same underlying unit. Because the cycles differ in length they drift out of alignment and return only after their least common multiple, so the composite pattern is far longer than either part. Distinguished from polyrhythm, where the cycle is shared and the subdivision differs.Demonstrated on the Rhythm Playground bonus page.
polyrhythm Polyrhythmik
Two or more parts dividing the same span of time into different numbers of equal parts — three against four being the textbook case, where the two lines coincide only at the start of each cycle. The ear can follow either division as the beat, and which one it hears is partly a matter of attention, which is what makes the effect so unstable and so useful. Distinguished from polymeter, where the pulse is shared and the cycle lengths differ.Demonstrated on the Rhythm Playground bonus page; related to Chapter 4 — Sound Authoring & Casting.
pan Panning (Panoramaregelung)
To position a mono source between the left and right channels of a stereo (or wider) output by adjusting their relative gains. Constant-power pan laws keep the perceived loudness uniform across the soundstage.Introduced in Chapter 9 — 3D Sound & Auralization.
pan control Panoramaregler
The knob, fader or automation lane on a mixer channel that sets the source’s position in the stereo (or multichannel) field. Implements a pan law — commonly −3 or −4.5 dB at center — to maintain loudness as a sound is swept.Introduced in Chapter 5 — Miking & Recording.
partial Teilton
Any single sinusoidal component of a complex sound. A partial may be harmonic, when its frequency is an integer multiple of the fundamental, or inharmonic when it is not.Introduced in Chapter 3 — Spectra, Modulation & Phase.
pass band Durchlassbereich (Passband)
The range of frequencies that a filter allows through with little or no attenuation, bounded by the cut-off frequencies. Together with the stop-band and transition slope it defines the filter’s response.Introduced in Chapter 7 — Effects & DSP.
peak indicator Peak-Anzeige (Spitzenpegelanzeige / Übersteuerungs-LED)
A meter element — typically an LED or numeric readout — that lights or holds whenever the input signal exceeds a chosen threshold, usually just below digital full-scale. Peak indicators complement average-reading VU meters by catching short transients.Introduced in Chapter 6 — Digitization & Editing.
peak value Spitzenwert
The largest instantaneous magnitude reached by a waveform during a given interval, measured from the zero baseline. Important for setting record levels so that transients do not clip.Introduced in Chapter 2 — Intensity & Loudness.
peak-to-peak Spitze-Spitze-Wert (Spitzen-Spitzen-Wert)
The instantaneous voltage or amplitude swing of a waveform from its lowest sample to its highest, expressed as a single number (Vpp or App). Distinct from dynamic range, which describes the usable span between the noise floor and the clipping ceiling of a system.Introduced in Chapter 2 — Intensity & Loudness.
peaks Spitzen (Pegelspitzen)
Brief, high-amplitude transients in a signal that rise well above its average (RMS) level — drum hits, plosives, attack edges. Adequate headroom is allotted above the nominal level to accommodate them without clipping.Introduced in Chapter 6 — Digitization & Editing.
Penrose staircase Penrose-Treppe
An impossible figure devised by Lionel and Roger Penrose in which a flight of stairs appears to ascend (or descend) in a closed loop. It is the visual analog of the auditory Shepard tone.Introduced in Shepard Tones.
periodic waveform periodische Wellenform
A waveform that repeats its shape exactly after a fixed interval, the period; its inverse is the frequency. Pure tones, square waves and the steady portion of a sustained instrument note are well approximated as periodic.Introduced in Chapter 1 — Communication, Frequency & Pitch.
phantom power Phantomspeisung
+48 V DC supplied along the same two balanced conductors of an XLR microphone cable that carry the audio signal, powering a condenser microphone’s internal head amplifier and, in externally polarized designs, the capsule-polarizing circuit as well. An electret capsule is permanently charged and draws only the head-amplifier supply. Standardized as IEC 61938; dynamic and most ribbon microphones can be left connected without harm. Vintage, damaged, or miswired ribbon microphones can however be damaged by phantom power; check the mic’s specification before patching it into a phantom-powered preamp.Introduced in Chapter 5 — Miking & Recording.
phase velocity Phasengeschwindigkeit
The speed at which a point of constant phase — a single crest — advances through space: vp = ω/k = fλ. It is a property of one frequency component rather than of a signal as a whole, so it need not match the group velocity at which a packet, and its energy, actually travels; where the two differ the medium is dispersive. A fourth sense of “phase”, alongside position within a cycle, absolute phase, and the relative phase between two signals.Demonstrated on the Phase & Group Velocity bonus page.
phase vocoding Phasenvocoder
A frequency-domain analysis-synthesis technique that uses a short-time Fourier transform (STFT) to track the amplitude and phase of a signal’s sinusoidal partials, allowing independent modification of time and pitch on resynthesis. The basis of high-quality time-stretching and pitch-shifting algorithms.Introduced in Chapter 7 — Effects & DSP.
phon Phon
The unit of loudness level. A sound is at n phons when listeners judge it as loud as a 1 kHz tone at n dB SPL, so the phon folds the ear’s frequency response into the decibel: the same 60 dB SPL is 60 phons at 1 kHz but considerably fewer at 50 Hz, where hearing is less sensitive. Each equal-loudness contour is a line of constant phons. Being a decibel measure it says how much louder one sound is than another only in the ordinal sense; for ratios, see sone.Introduced in Chapter 2 — Intensity & Loudness.
phonemes Phoneme
The smallest contrastive sound units of a spoken language — e.g. the /p/ and /b/ that distinguish pat from bat. Phonetic editing and concatenative speech synthesis operate at this level.Introduced in Chapter 6 — Digitization & Editing.
pink noise rosa Rauschen
A random signal whose power spectral density is inversely proportional to frequency, so that each octave contains equal energy. Pink noise sounds spectrally balanced to the ear and is the standard test signal for tuning sound systems and rooms.Introduced in Chapter 3 — Spectra, Modulation & Phase.
pinnae Ohrmuschel (pl. Ohrmuscheln)
The visible, cartilaginous outer ears (singular pinna). Their convolutions impose direction-dependent spectral notches and peaks on incoming sound, providing the principal cue for elevation and front/back disambiguation in the HRTF.Introduced in Chapter 9 — 3D Sound & Auralization.
pitch class Tonklasse (Pitch-Class)
The equivalence class of all pitches that share the same chroma — all the C’s in every octave belong to the pitch class C. Pitch class abstracts a note’s name from its register and underlies the circular representation of pitch.Introduced in Shepard Tones.
pitch shift Tonhöhenverschiebung (Pitch-Shift)
A processing operation that changes the perceived pitch of a sound by a chosen interval while ideally leaving its duration and timbre intact. Modern implementations use phase vocoding, sinusoidal modeling or PSOLA.Introduced in Chapter 7 — Effects & DSP.
pizzicato pizzicato (gezupft)
A string-playing technique in which the string is plucked with a finger instead of bowed, producing a sharply attacked, quickly decaying note rich in inharmonic onset content. Abbreviated pizz. in scores.Introduced in Chapter 3 — Spectra, Modulation & Phase.
playback Wiedergabe
The reproduction of a stored or transmitted audio signal through D-A conversion, amplification and a transducer (headphones or loudspeakers). Playback conditions — room, speakers, level — substantially shape what the listener hears.Introduced in Chapter 4 — Sound Authoring & Casting.
pre-amplification Vorverstärkung
The first gain stage in an audio chain, used to raise a microphone or instrument signal from millivolts up to line level so that subsequent processing and conversion stages operate at their nominal level with low noise.Introduced in Chapter 5 — Miking & Recording.
presbycusis Altersschwerhörigkeit (Presbyakusis)
Age-related hearing loss, primarily affecting the highest audible frequencies. Begins gradually in the third decade and accelerates after ~50, raising the threshold of audibility and shifting the high-frequency end of the equal-loudness contours upward. See the Fig 2.3 interactive demo (slide the Listener age slider).
proximity effect Nahbesprechungseffekt
The bass boost exhibited by directional (cardioid, bi-directional) microphones as the source moves close to the diaphragm, caused by a steeper pressure-gradient at short distances. Often exploited to add warmth or intimacy to spoken voice.Introduced in Chapter 5 — Miking & Recording.

Q

quantization Quantisierung
The mapping of each sampled amplitude to the nearest value of a finite set of representable codes, introducing a small rounding error called quantization noise. Word lengths of 16 bits give about 96 dB of dynamic range; 24 bits give about 144 dB.Introduced in Chapter 6 — Digitization & Editing.

R

R/D ratio R/D-Verhältnis (Verhältnis Hall- zu Direktschall)
The ratio of reverberant-field to direct-field sound energy at a listening position, R/D. It increases with distance from the source and with room reverberance, and is the primary cue for perceived source distance in a reverberant space.Introduced in Chapter 7 — Effects & DSP.
rarefaction Verdünnung (Unterdruckphase)
The half of a sound wave in which air molecules are momentarily more sparsely packed than in the surrounding medium — the negative-pressure counterpart of compression.Introduced in Chapter 1 — Communication, Frequency & Pitch.
ray tracing Strahlverfolgung (Ray-Tracing)
A geometric room-acoustics method that follows the paths of many sound “rays” emitted from a source as they reflect, scatter and absorb at boundaries, building up an impulse response for the listener position. Used in acoustical CAD and auralization.Introduced in Chapter 9 — 3D Sound & Auralization.
RCA jacks Cinch-Stecker (RCA)
Unbalanced single-conductor coaxial connectors widely used on consumer audio gear, color-coded white for left and red for right. Also known as phono plugs; not to be confused with the balanced TRS or XLR connectors used in professional work.Introduced in Chapter 6 — Digitization & Editing.
receiver Empfänger
The end of the audio transmission chain that takes in the produced sound — the listener’s ear, a microphone capturing playback, or any analogous sensor.Introduced in Chapter 1 — Communication, Frequency & Pitch.
recording Aufnahme (Tonaufnahme)
The capture of an audio signal in a storage medium for later playback or editing. Digital recording involves microphone or line pickup, preamplification, A-D conversion, and storage to disk or memory.Introduced in Chapter 4 — Sound Authoring & Casting.
relative dB relative dB (Pegel relativ zur Bezugsgröße)
A decibel measurement made against an arbitrary reference rather than an absolute physical unit. Levels on a DAW meter (dBFS), on a console (dBu, dBV) or in signal-to-noise figures are all relative-dB quantities.Introduced in Chapter 2 — Intensity & Loudness.
relative phase relative Phase (Phasendifferenz)
The phase difference, in degrees or radians, between two signals at the same frequency. Audible whenever two waveforms are summed acoustically or electrically: identical signals 180° out of phase cancel; in-phase signals reinforce.Introduced in Chapter 3 — Spectra, Modulation & Phase.
reverberation Nachhall
The persistence of sound in an enclosure after the direct source has stopped, produced by the superposition of countless reflections off walls, ceiling and floor. It conveys room size and character and substantially aids the externalization of binaural images.Introduced in Chapter 9 — 3D Sound & Auralization.
reverberation time Nachhallzeit
The time, in seconds, taken for the reverberant sound field to decay by 60 dB after the source ceases (RT60). It depends on room volume and total absorption, and is the most perceptually salient parameter of a reverberant space.Introduced in Chapter 7 — Effects & DSP.
room mode Raummode (Raumeigenmode)
A resonance of a room at a frequency whose wavelength matches a dimension (or a simple ratio of dimensions) of the enclosure, producing standing-wave patterns of peaks and nulls in the sound field. Low-frequency modes dominate small rooms and account for much of the ‘boomy in one corner, thin in another’ problem that bass traps and speaker placement aim to fix.Introduced in Chapter 8 — Playback & Sound Check.
room tone Raumton (Atmo, Grundgeräusch eines Raums)
The sound a location makes when nothing is happening in it: ventilation, a compressor cycling, traffic through glass, and the room’s own response to all of it, colored by its modes and reverberation. Recorded deliberately on location — thirty seconds to a minute at every setup, same microphone, same position, same gain as the dialogue — because an editor needs it to bed re-recorded dialogue, to patch the holes left where an unwanted noise was cut out, and above all to keep the background continuous across cuts. Silence will not serve: a gap of true digital silence reads as a technical fault rather than as quiet, the background falling away where the ear expects it to persist. Room tone is as particular as a fingerprint, so a bed borrowed from another location rarely passes. Sometimes called presence. It differs from ambience (also atmos, or a wild track), which is background carrying identifiable content that belongs to the storytelling — a playground, a harbor, a clock — and from the noise floor, which is the measured level of the residue rather than the sound of it.Introduced in Chapter 5 — Miking & Recording.
root mean square (RMS) value Effektivwert (RMS)
A measure of a waveform’s effective amplitude, obtained by squaring its instantaneous values, averaging over a chosen interval, and taking the square root. For a sine wave RMS equals peak/√2; RMS levels track perceived loudness much better than peaks.Introduced in Chapter 2 — Intensity & Loudness.

S

sample-rate conversion Abtastratenwandlung (Sample-Rate Conversion)
Resampling a digital signal from one sample rate to another — up-sampling (e.g. 22.05 kHz → 44.1 kHz, see up-sampling), down-sampling, or non-integer conversions such as 48 kHz → 44.1 kHz. High-quality resamplers use polyphase filters to avoid aliasing and to preserve transient detail.Introduced in Chapter 7 — Effects & DSP.
sampling Abtastung
The periodic measurement of an analog signal at uniformly spaced instants in time, producing the discrete-time sequence of values that a digital system stores and processes.Introduced in Chapter 6 — Digitization & Editing.
sampling rate Abtastrate
The number of samples per second taken from an analog signal by an A-D converter, measured in hertz. By the Nyquist theorem it must strictly exceed twice the highest frequency to be represented; standard audio rates include 44.1, 48, 88.2, 96 and 192 kHz.Introduced in Chapter 6 — Digitization & Editing.
Shepard tone Shepard-Ton
An auditory illusion devised by Roger Shepard, in which a stack of octave-spaced sinusoids weighted by a fixed bell-shaped spectral envelope glides upward (or downward) in pitch; partials cycle through the envelope so the tone appears to ascend (or descend) endlessly without ever changing register.Introduced in Chapter 9 — 3D Sound & Auralization.
shielded cable abgeschirmtes Kabel
An audio cable whose signal conductors are surrounded by a conductive shield (braid, foil or both) tied to ground, intercepting external electromagnetic interference before it can be induced into the signal. Standard for microphone and line-level interconnects.Introduced in Chapter 5 — Miking & Recording.
short-time Fourier transform (STFT) Kurzzeit-Fourier-Transformation (STFT)
A time-frequency analysis that takes successive short, overlapping windows of a signal, multiplies each by a smoothing window function (Hann, Hamming, …), and applies an FFT to each. The resulting sequence of complex spectra — one per analysis frame — tracks how the signal’s magnitude and phase content evolve over time, and is the basis for spectrograms, phase vocoding, and most real-time spectral effects.Defined here; not discussed in the chapter text.
sideband Seitenband
A spectral component created by modulation, appearing above and below the carrier at frequencies of carrier ± n × modulator, for integer n. In frequency modulation the sidebands form an evenly-spaced comb whose amplitudes follow the Bessel functions of the modulation index; it is the appearance of these sidebands that turns an audio-rate wobble into a timbre.Introduced in Chapter 3 — Spectra, Modulation & Phase; visualized on the Sirens & FM bonus page.
signal processing Signalverarbeitung
Any operation applied to an audio signal — filtering, equalization, dynamics, modulation, reverberation, encoding — whether realized in analog circuitry or in software.Introduced in Chapter 1 — Communication, Frequency & Pitch.
signal-quantization noise ratio (SQNR) Signal-Quantisierungsrausch-Verhältnis (SQNR)
The ratio, in decibels, between a signal’s level and the noise introduced by amplitude quantization. For an ideal uniform quantizer with N-bit words it is approximately 6.02·N + 1.76 dB, i.e. about 98 dB at 16 bits.Introduced in Chapter 6 — Digitization & Editing.
signed integer vorzeichenbehaftete Ganzzahl
An integer encoding that represents both positive and negative values around zero, almost always via the two’s-complement scheme — the bit pattern is read so that the most-significant bit contributes a negative weight. Standard for PCM audio (16- or 24-bit) where samples swing symmetrically about silence.Introduced in Chapter 2 — Intensity & Loudness.
sine wave Sinuswelle
The simplest periodic waveform, with a single frequency and no harmonic content; mathematically described by A sin(2πft + φ). Any complex periodic waveform can be expressed as a sum of sine waves (Fourier series).Introduced in Chapter 1 — Communication, Frequency & Pitch.
slope Filtersteilheit (Flankensteilheit)
The rate at which a filter’s magnitude response rolls off in its transition band, expressed in decibels per octave (or per decade). A first-order section has a 6 dB/octave slope; cascading sections increases the slope in 6 dB/octave steps.Introduced in Chapter 7 — Effects & DSP.
sound designer Sounddesigner (Sounddesignerin)
The author of the non-musical sonic content of a film, game or installation: dialogue editing, sound effects design, ambience, and the integration of music with picture. The boundary with the role of composer is fluid in contemporary practice.Introduced in Chapter 1 — Communication, Frequency & Pitch.
sound effect (SFX) Klangeffekt (Geräusch / SFX)
Any non-speech, non-music sound used to convey action, atmosphere or punctuation in a media work — footsteps, gunshots, doorslams, wind, UI beeps. Conventionally abbreviated SFX in script, edit-list and audio-post practice. Sound effects can be recorded on location, Foley-performed, drawn from libraries, or synthesized.Introduced in Chapter 4 — Sound Authoring & Casting.
sound file compression Audiodatei-Kompression (Tondateikompression)
Encoding an audio file so it occupies less storage or bandwidth. Lossless schemes (FLAC, ALAC) preserve every sample; lossy schemes (MP3, AAC, Opus) discard perceptually masked content to achieve much larger savings.Introduced in Chapter 6 — Digitization & Editing.
sound modification functions Klangbearbeitungsfunktionen (Sound-Modifikationen)
DSP operations that transform an audio signal’s content — level, dynamics, spectrum, time, pitch or spatial position. They form the creative palette of effects processing as distinct from utility operations such as normalization or sample-rate conversion.Introduced in Chapter 7 — Effects & DSP.
sound pressure level meter Schallpegelmesser (Schalldruckpegelmesser)
A calibrated instrument that measures sound pressure level in dB SPL, referenced to 20 µPa. Standard meters offer A-, C-, and Z-weighting curves and fast/slow time constants per IEC 61672.Introduced in Chapter 8 — Playback & Sound Check.
sound system Tonanlage
An assembly of electroacoustic equipment — sources, processors, amplifiers and loudspeakers (or headphones) — that delivers recorded or live audio to listeners.Introduced in Chapter 1 — Communication, Frequency & Pitch.
source (Schall-)Quelle
The originating element of an audio chain — an acoustic instrument, a voice, a microphone’s output, an oscillator, a synthesized signal, a recorded file.Introduced in Chapter 1 — Communication, Frequency & Pitch.
sone Sone
The unit of loudness itself, on a ratio scale: two sones is twice as loud as one sone, and four is twice as loud again. It is defined against the phon at 1 sone = 40 phons, with each doubling of sones costing about 10 further phons above that level. The distinction matters because the phon, being a decibel scale in disguise, cannot be used arithmetically — 80 phons is not twice 40 phons, whereas 2 sones really is twice 1 sone. Proposed by S. S. Stevens in 1936.Introduced in Chapter 2 — Intensity & Loudness.
spatial location räumliche Position
The apparent position of a sound source in space, characterized by direction (azimuth and elevation) and distance. Spatial location is conveyed by interaural time and level differences, spectral HRTF cues, and direct/reverberant balance.Introduced in Chapter 9 — 3D Sound & Auralization.
spectra
See spectrum.Introduced in Chapter 3 — Spectra, Modulation & Phase.
spectral balance spektrale Balance (Klangbalance / Frequenzbalance)
The relative distribution of energy across frequency bands in a signal. Equalization, microphone choice and room acoustics all shape spectral balance; a well-balanced mix sounds neither dull nor thin on a range of playback systems.Introduced in Chapter 3 — Spectra, Modulation & Phase.
spectrogram Spektrogramm
A picture of how a signal’s spectrum changes over time: time runs along one axis, frequency along the other, and the strength of each frequency component at each instant is rendered as brightness or intensity. Computed frame by frame from the short-time Fourier transform, it is the standard display for reading pitch glides, harmonic structure, formants, and modulation sidebands — the evenly-spaced sidebands of a frequency-modulated siren, for example, appear as a ladder of parallel horizontal traces.Introduced in Chapter 1 — Communication, Frequency & Pitch; also in Chapter 3 and the Chapter 7 waterfall; and on the Sirens & FM and Shepard-tone pages.
spectrum Spektrum
The distribution of a signal’s energy across frequency, obtained mathematically by the Fourier transform. The spectrum — both its instantaneous shape and its evolution in time — underlies the perception of timbre.Introduced in Chapter 3 — Spectra, Modulation & Phase.
SPL meter
See sound pressure level meter.Introduced in Chapter 8 — Playback & Sound Check.
splicing Schnitt (Tonschnitt / Splicing)
The joining of two audio segments at a chosen point. Originally a razor-blade cut of analog tape, it is now performed sample-accurately on a DAW timeline, usually with a short crossfade to avoid clicks or taps.Introduced in Chapter 6 — Digitization & Editing.
spot miking Stützmikrofonierung (Spot-Miking)
Placing a microphone close to an individual instrument or voice to isolate it within a larger ensemble pickup. The spot signal is then mixed with main or distant microphones to combine detail with overall blend.Introduced in Chapter 5 — Miking & Recording.
square wave Rechteckwelle
A periodic waveform that alternates between two fixed levels with equal duration in each state. Its spectrum contains only odd harmonics of the fundamental, with amplitudes proportional to 1/n.Introduced in Chapter 3 — Spectra, Modulation & Phase.
stereo miniature jack 3,5-mm-Stereoklinke (Mini-Klinke)
A three-conductor 3.5 mm TRS connector that carries stereo line-level audio on a single plug; ubiquitous on consumer headphones, laptops and portable players.Introduced in Chapter 6 — Digitization & Editing.
stop band Sperrbereich (Stoppband)
The range of frequencies that a filter attenuates strongly, beyond the cut-off and the transition band. Specified by the minimum attenuation maintained across the band (e.g. −60 dB).Introduced in Chapter 7 — Effects & DSP.
storage Speicher (Speichermedium)
The medium in which captured or processed audio is held for later retrieval — today, typically disk files in WAV, AIFF, FLAC or compressed formats, on local drives, network shares or cloud services.Introduced in Chapter 4 — Sound Authoring & Casting.
swing Swing (Shuffle-Feel)
Playing a pair of written-equal notes unequally, long then short, so the second arrives later than the page says. The amount is conventionally given as the point within the beat at which that second note lands: 50 % is straight (two even eighth notes), 66.7 % is the triplet or shuffle feel (the beat in three, notated as a quarter plus an eighth under a triplet bracket), and 75 % is the hard dotted feel (the beat in four, notated as a dotted eighth plus a sixteenth). Most swung playing sits between the first two rather than on them, and the amount typically relaxes as tempo rises. The percentage scale is the one Roger Linn built into the drum machines that made swing adjustable, carried on by the Akai MPC and Logic; several other sequencers count the same effect from zero instead.Adjustable, and drawn, on the Rhythm Playground bonus page.
sweet spot Sweet Spot (optimaler Hörplatz)
The listening position at which a stereo or multichannel sound system produces its intended spatial image. For two-channel stereo it lies at the apex of an equilateral triangle whose base is the line between the loudspeakers.Introduced in Chapter 8 — Playback & Sound Check.

T

timbre Klangfarbe
The perceptual quality that distinguishes two sounds of the same pitch, loudness and duration — what makes an oboe sound different from a clarinet on the same note. Timbre depends on spectral content, the time evolution of partials, and onset/decay behavior.Introduced in Chapter 3 — Spectra, Modulation & Phase.
timeline Timeline
In African and Afro-diasporic music, a short repeating pattern — usually struck on a bell, clave, or other cutting voice — that the rest of the ensemble orients itself against. It is not the beat, and it is not a drum part: it is the shared reference every other line is heard in relation to, which is why it is struck on something that cuts through the texture. Son clave, rumba clave, the bembé standard pattern, and tresillo are all timelines; the same idea is also called a key pattern or a bell pattern. Because a timeline is one line rather than a whole kit, loading one into every instrument at once would put the ensemble in unison — the opposite of its purpose.Explored on the Rhythm Playground bonus page; related to polyrhythm and clave.
tone control Klangregelung (Tonblende)
A simple equalization control, typically a single bass and treble knob (or a low- and high-shelf pair), that broadly adjusts spectral balance without offering separate band parameters.Introduced in Chapter 3 — Spectra, Modulation & Phase.
transducer (Schall-)Wandler
Any device that converts energy from one form to another. In audio: microphones (acoustic to electrical), loudspeakers and headphones (electrical to acoustic), phonograph cartridges (mechanical to electrical), tape heads, and so on.Introduced in Chapter 1 — Communication, Frequency & Pitch.
triangle wave Dreieckwelle
A periodic waveform whose value rises and falls linearly between two extremes. Its spectrum contains only odd harmonics, with amplitudes proportional to 1/n2, giving a softer timbre than a square wave.Introduced in Chapter 3 — Spectra, Modulation & Phase.

U

up-sampling Aufwärtsabtastung (Upsampling)
Resampling a digital signal to a higher sample rate — conceptually, inserting zero-valued samples between the originals and then low-pass filtering to interpolate them smoothly. Useful for moving between rates (e.g. 22.05 kHz → 44.1 kHz) and inside oversampling DSP processors. The inserted zeros are an algorithmic step, not audible silence.Introduced in Chapter 7 — Effects & DSP.

V

vibrato Vibrato
A periodic, small-amplitude modulation of a tone’s fundamental frequency, typically at 5–7 Hz and within about a semitone. Vibrato is a deliberate expressive device in singing and bowed-string playing and is distinct from amplitude tremolo.Introduced in Chapter 3 — Spectra, Modulation & Phase.
virtual hearing virtuelles Hören
Hearing of sound that originates not at an actual acoustic source but at a transducer (loudspeaker or headphone) driven by a stored or transmitted signal — the mediated counterpart of natural hearing.Introduced in Chapter 1 — Communication, Frequency & Pitch.
virtual pitch virtuelle Tonhöhe (Residualtonhöhe)
The pitch heard at a fundamental frequency that the signal does not actually contain. Partials at 800, 1000, and 1200 Hz are heard as a 200 Hz tone: the hearing system reports the fundamental the pattern implies rather than only the energy that arrived. Two families of explanation compete, and both still have support — pattern matching, in which the resolved partials are fitted against a harmonic template, and periodicity detection, in which the pitch follows from the repetition rate of the waveform itself (see autocorrelation). Virtual pitch and fundamental tracking name the same phenomenon from opposite sides — the percept and the mechanism.Introduced in Chapter 3 — Spectra, Modulation & Phase.
voice coil Schwingspule (Tauchspule)
The cylindrical winding of fine wire attached to a dynamic microphone’s diaphragm (or a loudspeaker’s cone) and suspended in a permanent magnetic field. Motion of the diaphragm in the field induces a voltage; conversely, current in the coil produces force on the cone.Introduced in Chapter 5 — Miking & Recording.

W

waveform Wellenform
The plot of a signal’s instantaneous value over time. For sound it represents the time history of acoustic pressure (or an electrical analog of it), and is the most direct visual representation of a signal in a DAW.Introduced in Chapter 1 — Communication, Frequency & Pitch.
waveform cycle Wellenformzyklus (eine Periode)
One complete repetition of a periodic waveform, from any phase point back to the same phase point. Its duration is the period; its reciprocal is the frequency.Introduced in Chapter 1 — Communication, Frequency & Pitch.
Web Audio API Web Audio API
A standard JavaScript interface for generating, processing, routing, and analyzing audio in the browser. It provides oscillators, filters, gain, panning, convolution, FFT analysis, and a low-latency graph that the Sonic demonstrations use to render sound in real time.Introduced in Chapter 0 — Welcome.
white noise weißes Rauschen
A random signal whose power spectral density is constant across frequency, so equal energy falls in every hertz of bandwidth. Heard as a bright hiss; the standard test signal for system bandwidth and impulse-response measurements.Introduced in Chapter 3 — Spectra, Modulation & Phase.

X

XLR connector XLR-Steckverbinder (Cannon-Stecker)
A three-pin locking connector standardized for balanced microphone- and line-level audio interconnects. Pin 1 is shield/ground, pins 2 and 3 carry the differential signal. Phantom power for condenser microphones is delivered on the same two signal pins; see balanced audio and phantom power.Introduced in Chapter 5 — Miking & Recording.

Z

zero crossing point Nulldurchgang
An instant at which a waveform passes through zero amplitude. Editing splices, loop boundaries and crossfades are best placed at zero crossings (with matching slope) to avoid the discontinuities that produce audible clicks or taps.Introduced in Chapter 6 — Digitization & Editing.