Sound Authoring and Casting
How to choose, categorize, and assign functions to sound sources for audio production for video, games, and interactive media — from narration and music to sound effects.
The Sound Authoring Process
In Chapters 4–6 we’ll examine the first part of the communication chain described in Chapter 1, where we author sound to create a particular effect:
- Casting different sounds according to function (Chapter 4);
- Recording, storage, and playback of sound (Chapter 5); and
- Digitization and editing methods for digital sound (Chapter 6).
Sound waveforms as described in Chapter 1 can now be seen in the context of recording and digitization of a specific sound to create a virtual sound experience. What the process involves is choosing from among the many different types of sounds sources available to the composer, and then subsequently storing and editing these sounds via analog-digital (A-D) and digital-analog (D-A) conversion. These are the basic functions of a digital audio workstation (DAW) for audio production in a DAW; i.e., the sound hardware and software present within the computer. Subsequent chapters examine more advanced methods for processing audio.
Casting
The compositional process involves capturing many different types of sounds, editing, processing, and assembly for a final product. But first we have to have something worthwhile to record. Casting involves finding sounds that are both appropriate and worthwhile for recording for use in a particular audio production, and then assigning a function to that sound; the sound can then be termed an audio cast member. What’s appropriate to cast is completely up to the composer; only the imagination should limit which sounds are used for a particular application. What’s more difficult sometimes is identifying sound sources that have something special about them —a certain sonic richness. Even more important is the difficulty in identifying sounds that will successfully transfer to the digital audio in a DAW medium. It’s somewhat akin to photography; one can be in the midst of a beautiful landscape, but only so many of the potential scenes will transfer successfully according to the composer’s intent. Furthermore, they need to be sounds that can be captured by a recording device with a minimum of inherent noise and distortion. It is tempting to use pre-recorded sounds, and indeed the use of “second hand” sampled sounds in the hands of an expert can yield incredibly satisfying results (as with many types of pop music.) But in many media productions the same types of music and effects recur over and over again, boring the listener instead of exciting them.
A composer needs to be especially careful that the sound is not overused, owned, or cliché or else the immediate reaction of some listeners will be exactly like when hearing a radio station that they don’t like —listeners are capable of making extremely fast judgments to “tune out” upon hearing certain sounds. Hence, it’s as important to become proficient at recording original sonic material as it is to know how to capture and manipulate the sounds made by other people.
There are many types of sound sources usually needing to be cast in a media-production environment. For our purposes, we’ll divide them into three different categories, seen at the left of Figure 4.1. The first are sounds made by humans, in particular, vocal music and the spoken voice. The second category is “sound sources of definite pitch,” which would include musical instrument sound sources, according to one’s personal definition of music: “traditional,” “non-traditional,” samples, dubs, acoustic, or electronic sounds produced by samplers, synthesizers, and so on. Usually, musical sounds have a definite pitch, except for certain percussion instruments. Music can be defined for present purposes as a layer of sound that is not directly connected to an action within a visual metaphor or object that is seen. The third category is for sounds mostly used for sound effects which sometimes but not always have an indefinite pitch. This would include those sounds that are directly connected to a visual action, metaphor or object.
Figure 4.1. Relationship between different classifications of sound sources, and their eventual application within an audio production.
The following example () is practically a science for the class of cinematic sound effects known as Foley sound. The number of sounds that are connected to visual objects are infinite: toasters, wind, automobiles, latches, etc.
Ultimately, the problem of categorizing sounds according to characteristics of the original sound source is that it is arbitrary; many sounds could fit into any three of these categories from moment to moment. For instance, non-traditional sound sources can easily be manipulated to sound like musical instruments, and both music and sound effects can be used to suggest a mood or attitude on the psychology of the listener. As an alternative, it is easier to categorize sounds into possible cast members according to their function. Figure 4.1 also shows how sound source characteristics map into Narration, Music and Sound Effect functions. Note also that the boundaries are not permanently fixed between these functional categories; a single sound source can act as any three of these cast types, even in combination.
In the discussion that follows, we’ll adopt some fairly narrow working definitions of narration, music and sound effect functions, using traditional examples. The main point is that these three functions are commonly utilized simultaneously in the process of composing all of the audio for a media production. But there’s nothing wrong with imagining less traditional alternatives, such pounding a garbage can lid to form an understandable, primitive narrative that’s also musical ().
Narration
Narration is defined here as the use of the spoken word for the purpose of presenting a story to a listener —a sequence of ideas that forms an image within the listener’s imagination. While we use a number of different media to “tell the story” in a media production, the power of the spoken voice can override other types of cues, since the subtle power of different types of narrative style is profound.
A narrator’s voice is more similar to an actor in a movie than to a written narrative in that we can immediately understand important information regarding the characterization of the speaker and the mood of the situation just through listening to the narrator’s delivery. A monologue or a dialogue in a play or movie script are only words on paper; the characterization that arises through the narrative performance. The manner in which characterization affects delivery is manifested in such things as timing, variation in pitch, volume, and other features inherent to the range of variation capable by a single person. Additionally there are features to the voice that are indicative or unique to a particular individual or a group of people (including stereotypes of particular people). We can guess the age of the person, whether or not they’re happy, sad, young, their language, dialect, and even cultural standing.
Most people use a range of characterizations in their everyday speech without being conscious of the fact. We use a certain type of character that directly affects one’s vocal delivery towards particular persons; one voice for people we’re intimate with, another for strangers, another with parents or family, and so on. Most of these characterizations are probably absorbed culturally from learning from others, but they are also absorbed from mass media, especially television. These characterizations also help define societal gender roles. Note how a certain “authoritative intimacy” is often found in the monologue used in personal care product advertisements aimed at women, but not for a cross-gender product such as toothpaste.
Listen to the following examples of different narration of the following simple text and imagine the different scenarios that would accompany this particular dialogue — “here it comes now.”
What are the sonic characteristics that allow us to differentiate between these different examples? Using the methods for describing sound outlined in Chapters 1–3, it’s possible to see how varying pitch, intensity, the type of frequency modulation, and the temporal sequence of the words helps us associate certain scenarios and situations with the narrative. We match these physical acoustical variations with cultural associations of fear, anger, or boredom.
Music
Some people think of music in terms of a recognizable “tune,” while to others any assemblage of sound or even silence (thanks to John Cage) can be viewed as music. As if tied to the old adage “I don’t know much about art but I know what I like,” many people are quite limited in their perceptions of what constitutes acceptable music. For commercial audio production for video, games, and interactive media, it’s often safer to aim for the lowest common denominator. It’s a whole art form to create musical tunes that sound vaguely familiar but not familiar enough to incur royalties for the use of the song. The rights to certain tunes can also be purchased, or use ones that are in the public domain. The danger with popular music usage is for something that is intended to have a lifetime longer than what is considered to be popular.
Traditionally one considers melody, harmony and rhythm as separate domains, with melody considered the most important. This is a western notion; for instance, in Indian classical music, the importance of rhythm is paramount (). Rhythm read as a cycle rather than a line is the subject of the Rhythm Playground, where the repeating patterns of West African, Afro-Cuban, and Balkan practice are drawn as necklaces of onsets around a circle and can be run against one another as polyrhythms.
One technique commonly used in modern media is the ostinato. In fact, it’s easy to create a simple ostinato by taking a snippet of sound and then repeating it ( to hear a snippet; to hear the sound looped a few times). This gets overdone since most modern media platforms allow continuous sound loops. Looping a sound saves storage space but it will become noticeably repetitious. Usually the function of an ostinato pattern is to set a mood without being noticed.
Complex ostinati are much more interesting. For instance, the sound examples used at the start of Chapter 1 were based on an ostinato. In this example a single ostinato pattern is heard, played on the lower notes of the guitar ( to hear the example). In this example, there are two ostinato patterns, one in the piano, one on the marimba ( to hear the example).
We’ve already shown how music can change the mood of a particular text, in Chapter 1. If we vary the type of music used with a particular narration, we can alter the emotional content as well. For instance, to listen to a “cute” setting of a little boy’s voice. But with this type of music instead — — the feeling is completely transformed.
Note that without music, we would need to have changed the style of the delivery used in the narrative to create a different effect. Through the combination of a narration with musical accompaniment, we can manipulate the “deeper meaning” of the dialogue.
Musical accompaniment in advertising is a good source of the psychological effects of harmony. For instance, which of the three chord change combinations inspires more confidence in a product:
The first example uses what is termed in musical harmonic theory as an “unresolved” harmonic sequence. The second example uses no particular harmonic sequence at all; instead, a consonant harmonic combination is followed by a much more dissonant harmonic combination. In the third example, the harmonic sequence is resolved, using the consonant (and sometimes sickeningly sweet) sound of major 7th chords.
Sound Effects (SFX)
Sound effects (“SFX” for short) are used within media applications for two purposes, as a function of the type of visual object or action they are associated with. First, there are SFX that respond to no visual object on the screen but instead to a physical action by the user, such as clicking or tapping the mouse. These create a sensation of interactivity due to the real-world familiarity of physical actions resulting in sonic feedback.
The following are some of the stereo sounds developed by Jay Boersma for mouse interaction system alerts. Bored of ordinary monaural system beeps, Jay developed the following sounds for various system alerts in order to exploit stereo playback.
Many programs for all types of platforms have methods for attaching different SFX to various computer alerts. While the above examples are very good, the fact that people frequently turn the volume off for many computer alerts or change them after a period of time points to the need for flexibility and a compositional approach.
The term “found sound” comes from the equivalent concept in art, the “found object.” Sound designers obtain many of their best SFX by gathering whatever they “find” during the collection phase of the creative process. This is a period where an artist collects materials without regard to the execution of a particular goal or project; it can be one of the most enjoyable and creative periods of the artistic process.
In that spirit, the first author had to kill an hour or two while waiting in a friend’s office for them to return from a lesson. Luckily, a good microphone and a portable digital recorder were on hand (it is amazing what a simple space such as an office can yield). Try guessing what the following sounds are; below, the sounds are identified. Note that the literal realism of the sounds will be affected by the quality of the playback system, in particular, the loudspeakers. Sounds such as these are “rich” for potential SFX, as will be heard.
Found Sound 1 (): This is a fire alarm bell that was hanging on the wall. Note that with a struck object, the timbre can vary radically as a function of the physical make-up of the striking device, and to a lesser degree, as a function of the degree of force used in striking the object. A familiar example is how a snare drum can immediately sound more “rock” or more “jazz” depending on whether or not the drummer uses sticks (drum beaters made of wood) or “brushes” (drum beaters made of thin wires). In this example, a coin was used to strike the metal. The actual fire alarm would never be heard in this way; a metal beater plays the sound repeatedly, never allowing the amplitude envelope to die out completely.
Found Sound 2 (): It was made by rubbing paper together. With good enough ears, and a good enough playback system, this is recognizable as the sound of money —specifically, US currency— being counted with two hands. Perhaps unexpectedly, bankers, professional gamblers and blind adults are quite good at distinguishing the sound of the “real thing” from mere paper. A finely developed sense of timbre is cultivated out of practical necessity, here, the timbre that results from the fiber content and size of money. A more obvious example is the sound of a coin rattling ().
Found Sound 3 (): This one is pretty easy; anyone who has ever had a desk job recognizes the sound of a length of transparent tape being pulled out of a dispenser and then cut with the dispenser’s teeth. But there’s more to this sound than first meets the ear. Note that the sound overall is comprised of three distinct components: first, the sound of the glue separating the length of the tape from the roll; second, the silence that is associated with orienting the tape up and over the teeth, once pulled out; and third, the quick tearing sound of the teeth cutting the tape. The second silent component is important to the realism of the effect; remember that one has to pull tape at an angle upwards away from the teeth to both extend its length and achieve enough of a distance to have adequate force to cause the tape to separate. It can be interesting and even fascinating to dyed-in-the-wool SFX designers how a detailed study of sound reveals the complexity of what might otherwise be considered a mundane and forgettable everyday action. But the power of understanding this is revealed on trying to create this sound incorrectly.
Found Sound 4 (): This is recognizably some sort of water sound, but its source is a surprise. This sound was produced by simply rocking a plastic water bottle back and forth with just a bit of water inside; and then playing back the sound at half speed. The resulting illusion is of a much larger body of water being displaced by a large object; a bit like the sound of a houseboat tied to a dock as someone steps aboard.
Found Sound 5 (): This is an example of slightly more complicated processing. The intended effect is the sound of a whip striking something. The sound source was, surprisingly, a page being turned in a book (actually, a music score, which has larger and thicker paper than most books). to listen to the original sound. For processing, the pitch was shifted 50 % lower, and then equalization was used to emphasize frequencies in the region between 2–6 kHz, yielding the slightly distorted “biting” sound as the whip makes contact.
Found Sound 6 (): This was a latch on a case for transporting electronic equipment. If we paste the sound of keys jingling in front of the sound of the latch opening, we can create a type of non-verbal sonic narrative — . We have someone searching through a set of keys followed by the successful opening of something that was locked. What’s missing is the sound of the key inside the latch itself, as it passes the lock tumblers (). Putting it all together via software editing, we get this: .
Casting for an interface rather than for a scene splits the problem in two, and the two halves have names. An auditory icon signifies by resembling: a crumpling sound for discarding a file, a camera shutter for a photograph taken, a latch for something closing. Nothing has to be learned, because the listener already knows what the world sounds like, and the found sounds above are exactly the raw material for them. An earcon signifies by convention: a short abstract motif — two rising notes for success, a descending pair for failure — which means nothing until it has been met a few times, and thereafter means it precisely. The pun is on icon and ear, and it is deliberate; the visual parallel is the difference between a picture of a wastebasket and a red octagon.
The trade-off runs the same way as in casting generally. Auditory icons are immediately legible and run out fast, since most abstractions have no characteristic sound — there is nothing that “synchronization complete” naturally sounds like. Earcons cost the listener a little learning and then scale without limit, and can be built in families, so that a shared rhythm marks a category while pitch or timbre distinguishes its members. Most systems of any size end up using both, and the cliché warning above applies with more force to the first kind: a found sound that is already carrying a meaning elsewhere will bring that meaning with it.