Realism in Music Reproduction

Realism in Music Reproduction
Commentary
Ralph Glasgal
www.ambiophonics.net
June 1999

Ever since 1881 when Clément Ader ran signals from ten (likely randomly) spaced pairs of telephone carbon microphones clustered on the stage of the Paris Opera via phone lines to single telephone receivers in the Palace of Industry that were listened to in pairs seemingly spaced more by accident than design, practitioners of the recording arts have been striving to reproduce a musical event taking place at one location and time at another location and time with as little loss in realism as possible. While judgments as to what sounds real and what doesn’t may vary from individual to individual, and there are many who religiously hold that realism is not the proper concern of audiophiles, such views of our hearing life should not be allowed to slow technical advances in the art of realistic auralization that listeners may embrace or disdain as they please. In this article we will review some past and recent developments in this field and indicate areas where more could readily be accomplished by recording engineers, manufacturers and audiophiles using existing tools.

. . . in many cases the home experience can now exceed a live event in acoustic quality

What Is Realism In Sound Reproduction?

In this article, realism in staged music sound reproduction will usually be understood to mean the generation of a sound field realistic enough to satisfy any normal ear-brain system that it is in the same space as the performers, that this is a space that could physically exist, and that the sound sources in this space are as full bodied and as easy to locate as in real life. Realism does not necessarily equate to accuracy or perfection. Achieving realism does not mean that one must slavishly recreate the exact space of a particular recording site. For instance, a recording made in Avery Fisher Hall but reproduced as if it were in Carnegie Hall is still realistic, even if inaccurate. It is doubtful that any home reproduction system will be able to outperform a live concert in a hall the caliber of Boston’s Symphony Hall, but in many cases the home experience can now exceed a live event in acoustic quality. For example, a recording of an opera made in a smallish studio, can easily be made to sound better at home, using the methods described below, than it did to most listeners at a crowded recording session. One can also argue that a home version of Symphony Hall, where one is apparently sitting tenth row center, is more involving that the live experience heard from a rear side seat in the balcony with obstructed visual and sonic view.

In a similar vein, realism does not mean perfection. If a full symphony orchestra is recorded in Carnegie Hall but played back as if it were in Carnegie Recital Hall, one may have achieved realism but certainly not perfection. Likewise, as long as localization is as effortless and as precise as in real life, the reproduced locations of discrete sound sources usually don’t have to be exactly in the same positions as at the recording site to meet the standards of realism discussed here. (Virtual Reality applications, by contrast, often require extreme accuracy but realism is not a consideration.) An example of this occurs if a recording site has a stage width of 120° but is played back on a stage that seems only 90° wide. What this really means in the context of realism is that the listener has moved back in the reproduced auditorium some fifteen rows, but either stage perspective can be legitimately real. Finally, being able to localize a stage sound source in a stereo, surround sound or Ambisonics system does not guarantee that such localization will sound real. For example, a soloist reproduced entirely via one loudspeaker is easy to localize but almost never sounds real.

Reality Is In The Ear Of The Behearer

While it is always risky to make comparisons between hearing and seeing, I will live dangerously for the moment. If from birth, one were only allowed to view the world via a small black and white TV screen, one could still localize the position of objects on the video screen and could probably function quite well. But those of us with normal sight would know how drab, or I would say unrealistic, such a restricted view of the world actually was. If we now added color to our subject’s video screen, the still grossly handicapped (by our standards) viewer would marvel at the previously unimaginable improvement. If we now provided stereoscopic video, our now much less handicapped viewer would wonder how he had ever functioned in the past without depth perception or how he could have regarded the earlier flat monoscopic color images as being realistic. Finally, the day would come when we removed the small video screens and for the first time our optical guinea pig was able to enjoy peripheral vision and the full resolution, contrast and brightness that the human eye is capable of and fully appreciate the miracle of unrestricted vision. The moral of all this is that only when all the visual sense parameters are provided for, can one enjoy true visual reality. At the present time there is no visual recording or display system that any human being could mistake for the real thing, but the IMAX system is a tantalizing foretaste of what might soon be possible.

One can only achieve realism if all the ear’s expectations are simultaneously satisfied.

Since most of us are quite familiar with what live music in an auditorium sounds like, we can sense unreality in reproduction quite readily. But in the context of audio reproduction, the progression toward realism is similar to the visual progression above. To make reproduced music sound fully realistic, the ears, like the eyes, must be stimulated in all the ways that the ear-brain system expects. Like the visual example, when we go from mono to stereo to matrix surround to Ambisonics to multi-channel discrete, etc.(listed in order of increasing accuracy, assuming that a new multi-channel method will actually emerge that can outperform Ambisonics or as discussed below Ambiophonics) we marvel at each improvement, but since we already know what real concert halls sound like, we soon realize that something is missing. What is usually missing is completeness and sonic consistency. One can only achieve realism if all the ear’s expectations are simultaneously satisfied. If we assume that we know exactly how all the mechanisms of the ear work, then we could conceivably come up with a sound recording and reproduction system that would be quite realistic. But if we take the position that we don’t know all the ear’s characteristics or that we don’t know how much they vary from one individual to another or that we don’t know the relative importance of the hearing mechanisms we do know about, then the only thing we can do, until a greater understanding dawns, is what Manfred Schroeder suggested over a quarter of a century ago, and deliver to the remote ears a realistic replica of what those same ears would have heard when and where the sound was originally generated.

Four Methods Used To Generate Reality At A Distance

Audio engineers have grappled with the problem of recreating sound fields since the time of Alexander Graham Bell. The classic Bell Labs theory suggests that a curtain, in front of a stage, with an infinite number of ordinary microphones driving a like curtain of remote loudspeakers can produce both an accurate and a realistic replica of a staged musical event and listeners could sit anywhere behind this curtain, move their heads and still hear a realistic sound field. Unfortunately, this method, even if it were economically feasible, fails on the first two counts with any finite number of speakers. Such a curtain can act like a lens and change the direction or focus of the sound waves that impinge on it. Like lightwaves, sound waves have a directional component that is easily lost in this arrangement either at the microphone, the speaker or both places. Thus each radiating loudspeaker in practice represents a new discrete sound source of uncontrolled directionality, communicating directly with both ears and therefore generating comb filter interference patterns and pinna directional distortion not present on the live side of the curtain.

Finally this curtain of loudspeakers does not radiate into a concert-hall size listening room and so one would have, say, an opera house stage attached to a listening room not even large enough to hold the elephants in Act 2 of Aida. This lack of opera-house ambience wouldn’t by itself make this reproduction system sound unreal, even if the rest of the field were somehow made accurate, but it certainly wouldn’t sound perfect. The use of speaker arrays (walls of hundreds of speakers) surrounding a relatively large listening area have been shown to be able to synthesize any sound field in a room with remarkable accuracy. But while this technique may be useful in sound amplification systems in halls, theaters or labs, application to the playback of even multi-channel recordings in the home seems doubtful except for the use of speaker arrays at the sides and rear or even overhead to deliver truly diffuse, reconstituted reverberant ambience to the home listener.

In general, multi-channel recording methods or matrix surround systems (Hafler, SQ, QS, UHJ, Dolby, 5.1,etc.) seem like exciting improvements when first heard by long deprived stereo music auditors, but in the end don’t sound real.

The Binaural Approach

A second more practical and often exciting approach is the binaural one. The idea is that, since we only have two ears, if we record exactly what a listener would hear at the entrance to each ear canal at the recording site and deliver these two signals, intact, to the remote listener’s ear canals then both accuracy and realism should be perfectly captured. This concept almost works and could conceivably be perfected, in the very near future, with the help of advanced computer programs, particularly for virtual reality applications involving headsets or near field speakers. The problem is that if a dummy head, complete with modeled ear pinnae and ear canal embedded microphones, is used to make the recording, then the listener must listen with in-the-ear-canal earphones because otherwise the listeners own pinnae would also process the sound and spoil the illusion.

The real conundrum, however, is that the dummy head does not match closely enough any particular human listeners head shape or external ear to avoid the internalization of the sound stage whereby one seems to have a full symphony orchestra (and all of Carnegie Hall) from ear to ear and from nose to nape. Internalization is the inevitable and only logical conclusion a brain can come to when confronted with a sound field not at all processed by the head or pinnae. For how else could a sound have avoided these structures unless it originated inside the skull? If one uses a dummy head without pinnae, then, to avoid internalization, one needs earphones that stand off from the head, say, to the front. But now the direction of ambient sound is incorrect and side localization is not fully accurate. IMAX is an example of this off the ear method, as supplemented with loudspeakers. There is also a circumaural earphone design that places a tiny speaker just over the notch in the lower front part of the ear so that many of the pinna resonances are still normally excited for frontal sounds. A similar ear speaker over the upper rear part of the ear can provide a similar pinna-friendly input for rear originating sounds. Unfortunately, headshape differences between the dummy head and the listeners head remain, and the dummy head should not have modeled pinnae if these earphones are to be used.

The fact that binaural sound via earphones runs into so many difficulties is a powerful indication that individual head shapes and outer ear convolutions are critically important to our ability to sense sonic reality.

Wavefield Synthesis

A third theoretical method of generating both an accurate and a realistic soundfield is to actually measure the intensity and the direction of motion of the rarefactions and compressions of all the impinging soundwaves at the single best listening position during a concert and then recreate this exact sound wave pattern at the home listening position upon playback. This method is the one expounded by the late Michael Gerzon starting in the early 70’s and embodied in the paradigm known asAmbisonics. In Ambisonics, (ignoring height components) a coincident microphone assembly, which is equivalent to three microphones occupying the same point in space, captures the complete representation of the pressure and directionality of all the sound rays at a single point at the recording site. In reproduction, speakers surrounding the listener, produce soundwaves that collectively converge at one point (the center of the listeners head) to form the same rarefactions and compressions, including their directional components, that were recorded.

In theory, if the reconstructed soundwave is correct in all respects at the center of the head (with the listeners head absent for the moment) then it will also be correct three and one half inches to the right or left of this point at the entrance to the ear canals with the head in place. The major advantage of this technique is that it can encompass front stage sounds, hall ambience and rear sounds equally, and that since it is recreating the original sound field (at least at this one point) it does not rely on the phantom image mechanism of Blumlein stereo. On the other hand Ambisonic theory is mute on the subject of how the sounds coming from the various loudspeakers are modified by the ear pinna and the head shape and how a decoder might compensate for these effects.

Thus the Ambisonic method is not easy to keep accurate at frequencies much over 2000 Hz and must and does rely on the apparent ability of the brain to ignore this lack of realistic high frequency pinna, head and waveform localization input and localize on the basis of the easier to reconstitute lower frequency waveforms alone. This would be fine if localization, by itself, equated to realism or we were only concerned with movie surround sound applications.

Other problems with basic Ambisonics include the fact that it requires at least three recorded channels (if we are concerned about quality) and therefore can do little for the vast library of existing recordings. Back on the technical problem side, one needs to have enough speakers around the listener to provide sufficient diversity in sound direction vectors to fabricate the waveform with exactitude and all these speakers positions, relative to the listener, must be precisely known to the Ambisonic decoder. Likewise the frequency, delay and directional responses of all the speakers must be known or closely controlled for best results and as in all other loudspeaker systems the effects of listening room reflections must also be taken into account, or better yet, eliminated.

As you might imagine, it is quite difficult, particularly as the frequency goes up, to insure that the size of the reconstructed field at the listening position is large enough to accommodate the head, all the normal motions of the head, the everyday errors in the listener’s position, and more than one listener. Those readers who have tried to use the Lexicon panorama mode, the Carver sonic hologram or the Polk SDA speaker system, all designed to correct the higher frequency parts of a simple stereo soundfield at the listener’s ear by acoustic cancellation will appreciate how difficult this sort of thing is to do in practice, even when only two speakers are involved.

In my opinion, however, the basic barrier to reality, via any single point waveform reconstruction method, like Ambisonics, is its present inability, as in the binaural case, to accommodate to the effects of the outer ear and the head itself on the shape of the waveform actually reaching the ear canal. For instance, if a wideband soundwave from a left front speaker is supposed to combine with a soundwave from a rear right speaker and a rear center speaker etc. then for those frequencies over say 2500 Hz the left ear pinna will modify the sound from each such speaker quite differently than expected by the equations of the decoder, with the result that the waveform will be altered in a way that is quite individual and essentially impossible for any practical decoder to control. The result is good low frequency localization but poor or non-existent pinna localization. Unfortunately, as documented below, mere localization, lacking consistency, as is unfortunately the case in stereo, surround sound or Ambisonics is no guarantor of realism. Indeed, if we must sacrifice a localization mechanism, let it be the lowest frequency one.

Finally, one can make a case that one can have glorious realism, even without any detailed front stage localization, as long as ambient localization is directionally correct (as anyone who has sat in the last row of the family circle in Carnegie Hall can attest to).

Ambiophonics

The fourth approach, that I am aware of, I have called Ambiophonics. Ambiophonics, which borrows a little from Binaural and still less from Ambisonics, assumes that there are more localization mechanisms than are dreamed of in the previous philosophies and strives to satisfy all of the mechanisms, as far as is possible. It also takes the psychoacoustic position that absolute binaural positional accuracy, as opposed to absolute realism, is not as vital and furthermore, that this reproduction technology need only be concerned with reproducing staged acoustical musical events, not movies or virtual reality. The advantage of focusing on just one aspect of sonic reality is that this reality is achievable today, is reasonable in cost, and is applicable to existing LPs and CDs.

One basic element in Ambiophonic theory is that it is best not to record rear and side concert-hall ambience or try to extract it later from a difference signal or recreate it via waveform reconstruction, but to synthesize the ambient part of the field using real stored concert hall data to generate ambience signals using the new generation of digital signal processors. The variety and accuracy of such synthesized ambient fields is limited only by the skill of programmers and data gatherers, and the speed and size of the computers used. Thus, in time, any wanted degree of concert hall design perfection could be achieved. A library of the worlds great halls may be used to fabricate the ambient field as has already been done with startling success in the JVC XP-A1010. The number of speakers needed for ambience generation does not need to exceed six or eight (although speaker walls would be optimum) and is comparable to Ambisonics or surround sound in this regard, but even more speakers could be used as this synthesis method is completely scaleable and the quality and location of these speakers is not critical.

Ambiophonics is usually less limited as to the number of listeners who can share the experience at the same time compared to most implementations of other methods using a similar number of speakers. Fortunately, two to five people can be accommodated by Ambiophonics in several of its practical incarnations.

The other basic tenet of Ambiophonics is similar to Ambisonics and that is to recreate at the listening position an exact replica of the original pressure soundwave. However, Ambiophonics does this by transporting the sound source, stage, and hall to the listening room rather than a point wavefront to the ears. In other words, Ambiophonics externalizes the binaural effect, using, as in the binaural case, just two recorded channels but with two front stage reproducing loudspeakers and eight or so ambience loudspeakers in place of earphones. Ambiophonics generates stage image widths up to 120° with an accuracy and realism that far exceeds that of any other 2 channel reproducing scheme. While it hardly seems to be necessary, the use of four channels and four main front loudspeakers can produce a full 180° stage image, (see below) but I doubt the expense would be worth it since I for one have never attended a live performance, and had a seat, where the music came from anything approaching a full half circle.

For reasons outlined below, Ambiophonic reproduction does require that the two main front speakers subtend an angle of only about 10° each side of the listening position so as not to generate the kind of pinna angle distortion for central sounds that phantom-image-stereo, Ambisonic or surround sound speaker placement almost always gives rise to. Ambiophonics also requires that a small lightweight, sound absorbing panel be placed on edge, centered in front of the listening position so as to prevent the left front speaker signal from reaching the right ear and vice versa. While there are electronic means to accomplish this end, (Carver, Lexicon, Cooper-Bauck-Harmon Intl.) and extra speaker means (Polk or easily home made) none of these work as well as a small inexpensive panel. The Ambiophonic listener is free to rotate his head, rock back and forth, and undulate from side to side, without image shift, just as in a concert hall. Most audio enthusiasts imagine that the use of the panel will be objectionable on aesthetic grounds. I certainly wish I could think of a less problematical way to accomplish the same end, but, at least in practice, one gets used to the panel very quickly and soon wonders why anyone listens without one. The new lightweight materials make it easy to store the panel between sessions or provide extras for multiple listeners.

The Stereo Dipole, AES Preprint 4463

For those unalterably opposed to using a panel, Ole Kirkeby, and Philip A. Nelson of The University of Southampton with Hareo Hamada of Tokyo Denki University have developed an electronic version of the panel. They have shown that the ideal speaker spacing for a crosstalk cancellation sytem be it mechanical or electronic is about 10 degrees. They refer to two speakers placed so close together as a “stereo dipole”. The electronic filters required to cancel crosstalk in this arrangement are somewhat easier to design and are more effective since at the narrower angle there is little diffraction around the head for the correction signals and so HRTF correction is not critical. Pinna angle distortion of the correction signals is also not a major factor and so the crosstalk cancellation can be allowed to operate over the full upper frequency range without restricting the size of the listening area or generating the audible phasiness effects that afflict electronic crosstalk cancellation schemes for widely spaced loudspeakers. A simple, low cost, lightweight panel will still remain the best choice for critical listeners.

Since Ambiophonics is a binaural based system, it does not provide the Blumlein loudspeaker crosstalk signal that furnishes the lowest frequency phase shift localization cues for recordings made with a coincident microphone. But to counterbalance this, Ambiophonics, or any crosstalk elimination idea, is more compatible than is standard stereo with the overwhelming majority of non-coincident microphone recording arrangements and the improvement in HF localization more than compensates for any loss in coincident mic LF localization. Furthermore, depending on its size and absorbency, the barrier (and even its electronic cousins) loses its effectiveness at low frequencies thus allowing some crosstalk and therefore amplifying LF phase cues for coincident microphone recordings. One can also move a little further back from the edge of the barrier or use a smaller panel when listening to coincident mic recordings.

As in all realistic systems, room treatment is essential for a good result and I have found that reducing the room reverberation time to less that .2 seconds works well in this context especially if used in combination with very directional, diffraction-free, point-source, front channel loudspeakers as once recommended by Malcolm Hawksford in another context.

Other Contrasts Between Ambiophonics and Ambisonics

The really fundamental difference between Ambisonics and Ambiophonics is that Ambisonics attempts to fabricate the exact compressions and rarefactions, including their intensities and directions at each ear canal, by summing the outputs of a given array of sound emitters whose drive signals must be derived by computation from three f (front), s (side), directional velocity microphone signals and the o omnidirectional pressure signal. (Since most readers will not be familiar with the mathematical symbols used in Ambisonics I will use o instead of w for the omnidirectional signal, f for the front-rear x signal and s for the left-right side y signal. We ignore the ‘h’eight (z) axis signal here.) In theory, there could be a playback computer that was fast enough to process such a three channel mic input with the accuracy needed to produce a perfect spherical wave front of say the fifth degree and the fourth order up to 15kHz. Each user would also have to load his personal pinna response curve into this computer to get the correct waveform at the entrance to the ear

Be the first to comment on: Realism in Music Reproduction

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

NanoFlo (82)Dynamique Audio (61)IKIGAI Audio (65)

Stereo Times Masthead

Publisher/Founder
Clement Perry

Editor
Dave Thomas

Senior Editors
Frank Alles, Mike Girardi, Russell Lichter, Terry London, Moreno Mitchell, Paul Szabady, Bill Wells, Mike Wright, and Stephen Yan,

Current Contributors
David Abramson, Tim Barrall, Dave Allison, Ron Cook, Lewis Dardick, John Hoffman, Dan Secula, Don Shaulis, Greg Simmons, Eric Teh, Greg Voth, Richard Willie, Ed Van Winkle, Rob Dockery, Richard Doron, and Daveed Turek

Site Management  Clement Perry

Ad Designer: Martin Perry