/

Talk

John Chowning

January 25, 2018

We had a good, solid talk with John Chowning, inventor of FM synthesis. Since its first musical use, FM has greatly expanded the musical possibilities of digital instruments. Its impact on every imaginable genre cannot be overstated, from contemporary classical music to dubstep. It was the synth sound of the 1980s, immortalized through the Yamaha DX7. John, however, does not refer to it as an invention. According to him, FM is a gift of nature that was just waiting to be discovered.

Now, understanding FM may seem daunting at first, as it incorporates some fundamental properties of math, music and acoustics. The beauty of John’s discovery is that (once properly used in a synthesizer) you don’t need to fully understand it. Just use your hands and ears to intuitively produce musical results that are pleasing, surprising, harmonic or inharmonic to your heart’s desire.

How would you explain FM synthesis to a child?

I would show the child how he or she might begin clapping two hands together, faster and faster and faster, them jump to the computer and show that we can make the claps even faster than the child is able to clap, and have the child listen to what happens. How the rate of claps changes from once per second, gradually through 8 times per second, to 16 times per second, all continuously increasing the rate until the child begins to hear a pitch.

At some point, I would say: "why don't you hum the pitch that you hear?" Now, I would do the same thing in reverse, you hear the pitch which the child has hummed, maybe something like 400 Hertz, which would be pretty close to G above middle C. Then I would reverse it, and as it slows, ask them to jump at in the moment they think they can clap that fast, and then slow down the computer-produced clap. We've established the fact, that when things happen at a certain rate, about 20-30 times per second, you no longer hear things as individual claps. You begin to hear things as tone quality (timbre) and then pitch.

Then I would do the same thing using the computer, with a sinusoid changing pitch — a vibrato. With a violin at hand, I would show what vibrato is — at the same pitch that the child hummed — and let my finger go up and down the fingerboard at an increasing rate. Again, jump to the computer with a sine wave at the hum pitch of 400Hz, with a vibrato depth increasing to ±40Hz at a rate of 1Hz. Then gradually increase the vibrato rate from 1Hz to 400Hz. As a last step I would gradually increase the vibrato depth to ±400Hz and we have caused the quality of the tone at 400Hz to change. All of a sudden we hear frequency modulation synthesis as a model of the original violin. That's one way of explaining it!

It's a phenomenon that has to do with the auditory system, and I think it's partially understood why it happens. It can be intuitively understood when we connect it to a real-life case, like vibrato in a musical instrument, which is a special case of frequency modulation. Once we've got the sinusoid modulating the carrier from 1 -400 Hertz, then we can change the distance up and down the keyboard, and show how the quality of the tone changes with deviation. That's basically how I would explain the properties of modulation rate and modulation depth to a child.

(I would also change the order in the demonstration, which is equally, if not more interesting — that is, first increasing the deviation of the 400 Hz sinusoid from ±0Hz to ±400Hz at a rate of 1Hz and then gradually increasing the rate from 1Hz to 400Hz.)

Can you re-visit and describe what it was like when you made the discovery?

With my first project, back in 1964, I was experimenting with spatialization of sound. In creating this spatial model I needed to generate tones that had internal dynamism — that changed in the course of the tone where the direct signal would distinguish itself from the reverberant signal. So changing pitch, as in portamento or vibrato, is probably the most effective ways to do this, because of the constructive/destructive interference of phases when that signal mixes with the reverberant signal — that complex relationship.

I began experimenting with tones, and became distracted from my original purpose. I wondered to myself what happens if I continued making the vibrato deeper and, of course, realized what I guess became known as a “chirp” (whistles a shrill rendition of the onset of a FM sound), I would have a sinusoid of extreme depth. With the vibrato of this sinusoid modulating the carrier at an increasing rate, I kept going, and finally I realized, I was not hearing in the time domain any longer, that is rate of change related to time, rather, I realized that using only two oscillators, one modulating the frequency of the other, I was hearing tones that seemed to be changing in quality or timbre, to have complex relationships between partials.

I did a number of quick examples of producing partials that were in the harmonic series, with a simple integer ratio between the vibrato rate and the center frequency (which was what I called them, rather than 'modulating' and 'carrier' frequencies), and inharmonic partials by non-integer ratios between the parameters.

I realized that when I transposed these, they were all well behaved, it was the same structural relationship. With a ratio of 1:2 of 100 Hz, when I transposed to 200 Hz, 200 to 400, 400 to 800 and so on I decided that these were predictable, and I knew I was on to something. It took me eight hours, working late at night at the Artificial Intelligence Lab at Stanford. I was alone, except for one young engineer, David Poole. I was astonished at what I was hearing — the different sounds that I could produce and so was he.

At that time, a similar sound could only be done with many, many sinusoidal oscillators, summing them up so they together form the partials of an inharmonic or harmonic spectrum, which would be an aperiodic or periodic waveform. David looked at a radio engineering text and applied the equation of simple radio frequency FM to what I’d done where the carrier frequency was in the audio domain.

I knew nothing about the synthesizers that were in use at that time, in 1964, although Moog and a few others had their first commercial analog synthesizer model out that year. My world was that of contemporary classical music, the classical tradition, and I was interested in electronic music. I saw the usefulness of FM synthesis. It was economically appealing, because it was a simple equation. You did need a computer, though, a powerful computer capable of processing and producing signals. Powerful for the day, that is to say! I knew this would be of musical interest, so over the next years, I pursued this.

It was a wonderful moment, of course. I realized this was a great gift of nature. I should point out that this was linear synthesis. It is not modulated in the logarithmic frequency space but in the linear frequency space.

When you made the discovery, did you realize its massive future impact on music, not to mention popular music, right away?

No. I had no foot in that world. As I said, my world was electronic music. I thought that the electric organ industry, which was a pretty big industry, might be interested. It seemed to me that they would be interested, but that was in 1971-72. I signed over my rights to Stanford University’s Office of Technology Licensing and let them do the marketing and, first of all, the patent search. They took on the responsibility, and tried to find if the organ industry would be interested in an application. The story is they contacted, first of all, these well-known companies, for example Hammond and Wurlitzer, in the United States. I would create examples and show them the code that I used to generate these examples. They agreed that the sound examples were interesting, but they had no idea about implementation, because they knew nothing about the digital domain. They all made analog organs in those days.

Stanford asked a business school student to scour the world for other organ manufacturers as well as the American manufacturers, and they discovered that Yamaha was the biggest industry manufacturer in the world. Its presence in the USA was rather small, but in Asia, especially Japan, they sold tens of thousands of organs per year. So they contacted Yamaha. They sent a young engineer, Kazukiyo Ishimura (who eventually became Managing Director of Yamaha, as did the engineer Hiro Kato who was also assigned to the first development phase of FM synthesis), who was already aware of the digital domain, and within ten minutes, he understood exactly what we had done, and what the nature of the implementation would be. They were already exploring digital domain control features for their analog organs and wave table synthesis for a then distant digital future.

You have to remember that in 1973, the scale of digital implementation was not very great, but they saw that the future was going to be such that the computing power in the coming years would be sufficient to run a real-time synthesis engine. So they invested about eight years of research before the DX7 was produced, and that of course was the great hit of the decade. The investment was great, and I want to point out that I don't call FM synthesis an invention; rather I call it a discovery. We received a patent for it, but the DX7, that was not really my work. That was the work of maybe a hundred really good Japanese engineers at Yamaha. I helped out during those years, followed their work closely and listened to their prototypes, and I consulted with their engineers on improvements and different kinds of tones and especially parameters set as a function of pitch height (pitch scaling), but it really was Yamaha, and the specific research of some really good engineers at Yamaha, that resulted in the DX7.

Do you think the time and environment was crucial to your discovery?

Yes I do. I'll tell you why: Let's say the computers I was working with had been powerful enough for me to do my experiments in real time. I'm not at all sure that I would have made the discovery! Because the condition under which I was working, on a time-share machine, a few seconds of sound might take me nearly two hours. So the time it took, perhaps specifically the time between experiments, I had to think. These were discrete times: I would generate a sample of a sound that was 20 Hz, with a modulating frequency of 20 Hz and a deviation of 100 Hz. Then I would wait. Then I would listen. Then I would increase to another. If I'd had continuous control, I think I probably would have missed it. I could have let the carrier sweep through frequencies that were way too high, and I would have missed the points where they converged to harmonic spectra. That being the case, the fact that I had to sit and wait and think, and listen, and then think about what I heard, "what will be the next step?" greatly enhanced my ability in realizing the discovery.

You mean the fact that you couldn't do experiments in real time was actually beneficial?

I think so. That's my working process, even today. I do discrete experiments, then wait and think. It's not that I don't make use of the real-time controls, but it's a practice that's had huge pay-off over the years. So I think these were special circumstances that very much enhanced the likelihood of discovery.

In your research, and in your composition, the crossover between music and mathematics is remarkable. Did 1960s writers like Guy Murchie, who wrote about Pythagoras, irrational numbers and the intimate relationship between math and music, influence you at all?

I was certainly influenced by Euclid's work on the golden section, when I made these early tones. As you know there are many, many tones that can produce harmonic spectra, and many more that can produce inharmonic spectra. I looked for numbers that produced partials that were not close to harmonic partials.

π is not a very good one, because it is so close to 3:1. The √2, on the other hand, is a good one because basically it's a tritone, and it's not confused with an out-of-tune harmonic spectrum. When I looked for other ratios, I found that _e_ (Euler's number) is not good, but the golden ratio (Φ), 1.618... is a rich one.

I became interested in the golden ratio as a way of producing inharmonic partials with FM. I explored that domain, and how the Fibonacci sequence is related to the golden ratio, and at some point I realized that by producing spectra, with a carrier/modulator ratio of powers of the golden ratio, four of the low order side-band components were also powers of the golden ratio.

That was kind of a Eureka moment, when I was in Berlin in 1974-75. Now, to a mathematician that may be obvious, but it certainly was not to me! Besides, mathematicians aren't usually interested in FM tones.

I realized I could create music where the pitch space and the spectral space would both be based upon powers of the golden ratio. A piece that I began working on in Berlin in 1974, and finished in 1977, turned out to be a very successful piece. The name of the piece is Stria (read the text of the presentation of Stria by Pierre Boulez in 1980 at the Théâtre d'Orsay).

The musical exploration of irrational numbers has been a fascinating and unique domain of FM synthesis, and I also made use of that in a piece that I finished in 2011, Voices, a real-time piece written in Max/MSP for solo soprano and computer synthesis. The voice signal was used to trigger the events and the computer would produce the processing and spatialization of the voice in real time, with a quad speaker system. It had the same theoretical underpinning, based upon the golden ratio, but with a very different auditory surface.

What was it like being tutored by the great Nadia Boulanger in Paris?

Studying with Nadia Boulanger was like being a child, although I was 26 or 27 years old! It was very rigorous and demanding. The way she taught harmony and composition, for example. She taught me to hear things in a different way, everything as line. But probably the greatest gift was her enthusiasm for music. The pieces that we studied at a weekly class of all of her 25-30 students were always rich experiences.

While I was there, in Paris for three years, I had the opportunity of hearing new music. I heard Luciano Berio and Karlheinz Stockhausen. This was music that was quite different from the music that Nadia Boulanger cared for, and I found it very exciting! My interest in composing music for loudspeakers began at the Domaine Musical concerts produced by Pierre Boulez.

The idea of creating music with computers, a program for the spatialization of sound, was directly inspired by Stockhausen, especially his piece Kontakte (1958-1960). He set up four microphones around a rotating loudspeaker in such a way that he could produce circular motion of sound.

My piece Turenas, from 1972, demonstrated the power of computers to do more than that, with four speakers. I realized that the future of electronic music was going to be in the digital domain and not in the analog domain. A fact that György Ligeti also realized, and saw in the same way, when he was at Stanford in 1972. He was present at the first performance of Turenas. It was he who helped me the most, when he returned to Europe, where he told Boulez about what was going on with computers and music in California.

That's how I became involved, together with Jean-Claude Risset, in the formative ideas of how the digital domain would be present in the initial IRCAM configuration of departments. So, my path was set there, based on my initial experience with electronic music in Paris in the late 1950s, early 1960s while I was studying with Nadia Boulanger.

Do you have a favorite pop song from the first golden age of FM?

That's not really my world, but I did work with Toto), which was a group that recorded a pop song called Africa in 1981, using a precursor to the DX7 called GS1. It had eight operators, four FM pairs. I helped programming, together with Gary Leuenberger and David Bristow who also worked with Yamaha doing sound design, factory voices that would be simple to use for keyboard players. They were super keyboard players with super ears. The presets were made very much according to the touch and feel of that keyboard, with velocity control and after touch. The musical sensibilities I wouldn't know about, because I'm not a keyboard player. Africa was one of the songs I knew!

Another one I remember, three or four years later, was a very exuberant performance by a group called Casiopea. The piece was called Eyes of the Mind, and they used the DX7.

What is extraordinary is the number of chips used for the FM synthesis. In the GS1, it was something like 40 or 50 chips. Three years later, in the DX7, I think it was three or four! The advance of technology was amazing.

As for other pieces, people point out to me the use of FM in familiar tunes, but as I'm not so interested in popular music I can't tell you what they are. I know that dubstep in today's music uses FM. When I listen to the radio, I hear FM used all the time. I just don't know the artist.

What words of advice can you give to other bold sound explorers?

Well, any music that uses a sinusoid as an element of synthesis or music production could just as well use at least a simple FM pair. It is just two parameters that you can increase and decrease to change the nature of whatever sound is being produced. Even a wave table might as well become a modulated wave table, because of the surprises that happen. One of the wonders of FM synthesis is that it can be so surprising.

I noticed, in your instrument, that you put some boundaries on the possibilities so that one doesn't end up in a daze without understanding how you got there, or end up in silence.

Anyway, about the future, and the possible spectra of differentiated tones with frequency modulation synthesis, with all the possible configurations, there's cascade modulation, parallel modulation, many modulator modulation and many carrier modulation. All of these must be an infinite (frequency domain) space, so there's much to explore.

I didn't mention, in the course of relating the experience of developing FM, one of the breakthroughs was my simulation, using FM synthesis, of the female singing voice. That was a big event, made in the domain of FM synthesis as possible use for a product. The examples that I produced pushed Yamaha over the edge towards FM. They were considering other possible implementations of a digital instrument. The FM voice convinced them. 
 
The FM voice was based on the fact that I had used two or three carrier frequencies, each modulated by a modulator, the frequency of which was always the pitch frequency. The carrier frequencies would be the harmonic frequencies closest to the frequency of the desired formant for a given vowel, a frequency of, say, 1200 Hz, when I was generating a pitch of around, say, G4, which would be about 400 Hz, and a carrier frequency with a ratio of 3:1. The side-band components would then be at 1200-400, 1200-800... and the second lowest sideband frequency would be at the pitch frequency. The upper side-band components would be 1200+400, 1200+800 and so on. With the components in that series it sounded like a resonance when the modulation index at about 1.0. That's how I produced the FM singing voice formants (or resonances).

It was very successful, and I created a table during the experiments, based on the data points of the octaves, the half-octaves of two- three octaves, and then interpolated the values for indexing the formant frequencies, based upon Johan Sundberg’s data having to do with formant frequencies of the female singing voice. These were not so obvious as is the case of the male voice, and Johan happened to be at IRCAM at the time I was doing this.
 
He was very surprised that I could produce such rich voices; they were more realistic than his, because his synthesis was based on subtractive synthesis, and he was using analog synthesis techniques where filters could not reduce a complex wave to be near sinusoidal — as in a soprano’s “oo” vowel. It was a big breakthrough. That was in 1978, and I sent the examples to Yamaha who immediately sent an engineer over to meet with me. They wanted to know how I did this. That's when they decided to commit their engineering effort to FM.

I don't think they ever produced an instrument that made use of that. I'm not sure why. I tried to encourage them to do that! I hope they will produce one that makes use of this kind of representation of FM, based upon resonances. Musical instruments have strong resonances, so this technique, although it is best exemplified in the female singing voice, because it's so special and instantly identifiable because we all know it so well from birth — we all have mothers — but the same idea, creating strongly resonant instruments based upon multiple carrier frequencies, where the modulating frequency is always the pitch frequency, and the carrier frequencies are always at the harmonic closest to the resonance, is a great way to produce resonant tones. That's basically it.

Maybe you can think about doing that?

What's your impression of the advancement of computer processing, in terms of being able to do real-time FM synthesis? Has the progress been faster, or slower, than you had expected when you started working with computers in the 1960s?

I didn't expect it to be as fast as it is, I don't think anybody did. I'm speaking to you on a mobile phone, which probably has more power than all the computers that I ever used in those early years. It's unimaginable. The power of my laptop, in another lifetime I can't imagine ever exhausting its capabilities. It's astonishing. I once did a cost of computing comparison, based on the cost of memory for computers when I began, which was the most costly component of the computer in 1965 compared to 2005. The power of the laptop that I was using in 2005, that kind of memory would have cost something like 80 billion dollars in 1965!

It's unimaginable. The amazing thing today, the great thing, is the many centers dedicated to open source. You probably even have some people at Elektron using open source code! Kids today, for a couple of hundred dollars, can get a laptop, access the Internet and take over the world with their imagination, with the all-available tutorials and with all the public domain software out there! Just what's freely available at CCRMA, Stanford is amazing: documents, tutorials and education. A kid in India could completely educate herself just using what is available on the Internet at no cost.

Speaking of public domain, I'm very impressed with the FM synthesis patent from 1974. It describes so clearly exactly the method, and even a transcript of the computer program, everything you need in order to implement the technique. It truly works in the way patents are originally intended to work: once the exclusivity expires 20 years after the application, it is truly a gift to the world. It's in the public domain.

There were lots of complaints, really, on the fact that Stanford was patenting the technology and keeping it from other people, but that was not really the case at all. It's making it available to people. Someone who wants to use it commercially, they had to pay Stanford for that right, but the composers in all the non-commercial institutions, they had unlimited use of the technology.

It really freed up the technique. Otherwise, people keep these things secret and don't tell anyone else about them as they develop it. Having the patent meant that we could be completely public, and Stanford was an institution that had the resources and cache, let's say, to be able to protect the patent. There were many attempts to copy it or overcome it, and lots of reverse engineering.

I was particularly intrigued by the practical details, like the fact that when the modulator/carrier ratio is 1/√2, it is particularly useful for bell or trumpet-like sounds, depending on the setup; dynamically going from a spectra of complex to simple partials during the attack phase in the bell case, or from simple to complex in the case of the trumpet. Harmonically, perhaps, you could say the bell is a reverse trumpet!

Exactly.

The √2 is also an irrational number. You already gave a thorough answer on the subject in one of my previous questions to you, but I just wanted to hear your idea of why you think they have such an important musical role to play. Why does the same number that drove the Pythagoreans crazy 2500 years ago make a nice digital bell sound today?

I don't know. I used it because it's a number that we all know. It's in the mind, and it's easy to remember. The partials are far enough from one another so that you don't confuse it with the harmonic spectrum where the harmonics are mistuned.

Other "magic numbers", let's say, in my mind are π, but because it's so close to three, if you form a ratio of one to π, it sounds like it should be one to three, but it doesn't sound right. The interesting thing about the √2 is that the partials are far enough apart. The spacing of the partials is such that you don't confuse it with a mistuned harmonic spectrum, and so is the golden ratio. Those sideband components, when they reflect around zero, they too are far enough from the upper sideband components that you don't confuse it. It sounds interestingly inharmonic, and convincingly inharmonic as bells and gongs.

What's the best hardware implementation of FM synthesis so far, in your opinion?

I'm not so familiar with that world, but I do know that Dave Smith Instruments has implemented an FM patch in the Prophet 12, and I think that's a fairly reputable instrument, is it not?

It's an excellent instrument.

They do it a bit differently. All the frequencies, I believe, the carrier-modulating frequencies in the implementation, are pitch frequencies, so you don't get quite the same exactness that you get if you use linear frequencies. Let's say the carrier/modulator ratio is one to three. If the pitch frequency is a hundred Hertz, the modulating frequency is not three hundred Hertz; it would be the octave and the fifth above pitch. That gives a slightly different kind of modulation. I think that's how they do it. I don't know, are there any other keyboards? I don't know that world, basically, about what's been done with keyboard synths these days.

What you've done (with the Digitone), from what I can tell, looks like a rich palette of possibilities. There are fine guided decisions that have to be made in FM synthesis, and this idea of letting the FM units be sound generators and then processed in normal ways that synths process sound sources seems perfectly reasonable. It eliminates the part where you have to know anything about the theory of FM.

That's the idea, making FM real-time, hands-on and musically useful. That, I guess, was your intention when you made the discovery, too! Reading your original description makes me realize the sheer power of the method: what enormous musical potential and power of expression you can have with just eight parameters.

It seems to me that the metaphor for the instrument (Digitone) is to give the user kind of a ball of re-formable plastic that, with the FM, can be pulled and stretched in many different ways. Of all the hundreds of sub-sets of what FM can do, yours seems to be a very useful sub-set! It leaves out many of the sub-sets but lets the user intuitively explore this re-formable, shapeable ball of stuff, then put that through the normal processes of synthesis that we know and love, band-pass filters and low-pass filters and so on.

It's a nice metaphor! Back in the 1970's, were you in cahoots with the Californians next door who made the first computers for the masses, the Commodore 64 and Apple II?

No, but I was at Stanford with David Zicarelli, who was a graduate student there. He did the first Apple voicing program for Opcode Systems, and before that he also did the Editor/Librarian for the DX7. So I was close to Opcode Systems, because they were right here, in Palo Alto.

That was a big advance, that voicing program that David produced (that evolved into the first MIDI sequencer for the Apple Macintosh in 1986, MIDIMAC, then Opcode Sequencer and ultimately the Vision Sequencer). You could see all the parameters on the screen, rather than through the DX7 window, where you had to remember lots of stuff. Other than that, the only sustained contact I had with the hardware industry was Yamaha, really.

I'd like to end this interview on a philosophical note. What is the basic grain of the universe?

Well, I try to follow the physicists and what they think. I find it far beyond my ability to really comprehend. They give us nice metaphors to think about it, like black holes and such things, but it's heavy math, and I think it's hard for them to explain to us the quantum physics and photons, mysteriously clouded in certain circumstances: I don't know if they fully understand it either. If you think about these things enough, you just assume it's true, and I remember when I first began work with computers, I'd try to understand the basics of all this. My interest in computers began when I read Max Mathews's article (The Digital Computer as a Musical Instrument; Science, Vol. 142, 1963), so I would ask Dave Poole, who was then an undergraduate in applied mathematics and a hacker - in those days hacker had a positive connotation - and he taught me everything that I had to know to get started.

I read the literature, I knew what a bit is and what the energy of an electron is, but when I asked David: "Tell me what an electron is!" I realized he had no answer to that. No one has ever seen an electron. All we've seen is a representation of something that is apparently there because the consequences of what we see can't be explained without assuming there's an electron! I realized right then and there that I just have to accept that there are some things that I won't be able to understand, some of the basic stuff, the physical stuff of which the universe is made.

If you had asked me, instead, "what is the basic grain of life?" I would have given a very different answer, because that is something that has to do with experience. It has in it the essentials of being and of relationships. I'm not a God believer, but I believe that there's a lot more than what I can think of or conjure up, something greater than myself. That includes all of human experience, which is mediated through history by writings and by music by art. All of this is available to me. I'm able to live in wonder as to what exists, without invoking the idea of some fatherly figure. I can make decisions that have real explanations. Even from childhood, I asked: "Mom, if there's a God, who created God?" This kind of conundrum, I just dispensed with. To be able to live in wonder because all of human thought and experience and expression are passed on into the future, to all humans to come, a future which I think, in general, is getting better.

Thanks so much for this interview.

It's nice to talk to someone who's interested in my work and knows what I'm talking about!

Interview by Daniel Sterner
Photo by Gemma Plannel

More Kulture