PubMed Health⌕ Search

Biomedical subjects

Brian D Simpson

Publications and source records attributed to Brian D Simpson.

15 recordsLinked to original sources

Monaural speech segregation using synthetic speech signals.

When listening to natural speech, listeners are fairly adept at using cues such as pitch, vocal tract length, prosody, and level differences to extract a target speech signal from an interfering speech masker. However, little is known about the cues that listeners might use to segregate synthetic speech signals that retain the intelligibility characteristics of speech but lack many of the features that listeners normally use to segregate competing talkers. In this experiment, intelligibility was measured in a diotic listening task that required the segregation of two simultaneously presented synthetic sentences. Three types of synthetic signals were created: (1) sine-wave speech (SWS); (2) modulated noise-band speech (MNB); and (3) modulated sine-band speech (MSB). The listeners performed worse for all three types of synthetic signals than they did with natural speech signals, particularly at low signal-to-noise ratio (SNR) values. Of the three synthetic signals, the results indicate that SWS signals preserve more of the voice characteristics used for speech segregation than MNB and MSB signals. These findings have implications for cochlear implant users, who rely on signals very similar to MNB speech and thus are likely to have difficulty understanding speech in cocktail-party listening environments.

Adult↗

Isolating the energetic component of speech-on-speech masking with ideal time-frequency segregation.

When a target speech signal is obscured by an interfering speech wave form, comprehension of the target message depends both on the successful detection of the energy from the target speech wave form and on the successful extraction and recognition of the spectro-temporal energy pattern of the target out of a background of acoustically similar masker sounds. This study attempted to isolate the effects that energetic masking, defined as the loss of detectable target information due to the spectral overlap of the target and masking signals, has on multitalker speech perception. This was achieved through the use of ideal time-frequency binary masks that retained those spectro-temporal regions of the acoustic mixture that were dominated by the target speech but eliminated those regions that were dominated by the interfering speech. The results suggest that energetic masking plays a relatively small role in the overall masking that occurs when speech is masked by interfering speech but a much more significant role when speech is masked by interfering noise.

Adolescent↗

Across-ear interference from parametrically degraded synthetic speech signals in a dichotic cocktail-party listening task.

Recent results have shown that listeners attending to the quieter of two speech signals in one ear (the target ear) are highly susceptible to interference from normal or time-reversed speech signals presented in the unattended ear. However, speech-shaped noise signals have little impact on the segregation of speech in the opposite ear. This suggests that there is a fundamental difference between the across-ear interference effects of speech and nonspeech signals. In this experiment, the intelligibility and contralateral-ear masking characteristics of three synthetic speech signals with parametrically adjustable speech-like properties were examined: (1) a modulated noise-band (MNB) speech signal composed of fixed-frequency bands of envelope-modulated noise; (2) a modulated sine-band (MSB) speech signal composed of fixed-frequency amplitude-modulated sinewaves; and (3) a "sinewave speech" signal composed of sine waves tracking the first four formants of speech. In all three cases, a systematic decrease in performance in the two-talker target-ear listening task was found as the number of bands in the contralateral speech-like masker increased. These results suggest that speech-like fluctuations in the spectral envelope of a signal play an important role in determining the amount of across-ear interference that a signal will produce in a dichotic cocktail-party listening task.

Adult↗

Precedence-based speech segregation in a virtual auditory environment.

When a masking sound is spatially separated from a target speech signal, substantial releases from masking typically occur both for speech and noise maskers. However, when a delayed copy of the masker is also presented at the location of the target speech (a condition that has been referred to as the front target, right-front masker or F-RF configuration), the advantages of spatial separation vanish for noise maskers but remain substantial for speech maskers. This effect has been attributed to precedence, which introduces an apparent spatial separation between the target and masker in the F-RF configuration that helps the listener to segregate the target from a masking voice but not from a masking noise. In this study, virtual synthesis techniques were used to examine variations of the F-RF configuration in an attempt to more fully understand the stimulus parameters that influence the release from masking obtained in that condition. The results show that the release from speech-on-speech masking caused by the addition of the delayed copy of the masker is robust across a wide variety of source locations, masker locations, and masker delay values. This suggests that the speech unmasking that occurs in the F-RF configuration is not dependent on any single perceptual cue and may indicate that F-RF speech segregation is only partially based on the apparent left-right location of the RF masker.

Acoustic Stimulation↗

Interference from audio distracters during speechreading.

Although many audio-visual speech experiments have focused on situations where the presence of an incongruent visual speech signal influences the perceived utterance heard by an observer, there are also documented examples of a related effect in which the presence of an incongruent audio speech signal influences the perceived utterance seen by an observer. This study examined the effects that different distracting audio signals had on performance in a color and number keyword speechreading task. When the distracting sound was noise, time-reversed speech, or continuous speech, it had no effect on speechreading. However, when the distracting audio signal consisted of speech that started at the same time as the visual stimulus, speechreading performance was substantially degraded. This degradation did not depend on the semantic similarity between the target and masker speech, but it was substantially reduced when the onset of the audio speech was shifted relative to that of the visual stimulus. Overall, these results suggest that visual speech perception is impaired by the presence of a simultaneous mismatched audio speech signal, but that other types of audio distracters have little effect on speechreading performance.

Acoustic Stimulation↗

The impact of hearing protection on sound localization and orienting behavior.

The effect of hearing protection devices (HPDs) on sound localization was examined in the context of an auditory-cued visual search task. Participants were required to locate and identify a visual target in a field of 5, 20, or 50 visual distractors randomly distributed on the interior surface of a sphere. Four HPD conditions were examined: earplugs, earmuffs, both earplugs and earmuffs simultaneously (double hearing protection), and no hearing protection. In addition, there was a control condition in which no auditory cue was provided. A repeated measures analysis of variance revealed significant main effects of HPD for both search time and head motion data (p < .05), indicating that the degree to which localization is disrupted by HPDs varies with the type of device worn. When both earplugs and earmuffs are worn simultaneously, search times and head motion are more similar to those found when no auditory cue is provided than when either earplugs or earmuffs alone are worn, suggesting that sound localization cues are so severely disrupted by double hearing protection the listener can recover little or no information regarding the direction of sound source origin. Potential applications of this research include high-noise military, aerospace, and industrial settings in which HPDs are necessary but wearing double protection may compromise safety and/or performance.

Adolescent↗

Audio and visual cues in a two-talker divided attention speech-monitoring task.

Although audiovisual (AV) cues are known to improve speech intelligibility in difficult listening environments, little is known about their role in divided attention tasks that require listeners to monitor multiple talkers at the same time. In this experiment, a call-sign-based multitalker listening test was used to evaluate performance in two-talker AV configurations that combined zero, one, or two channels of visual information (neither, one, or both talkers visible) with zero, one, or two channels of audio information (no audio, both talkers played from the same loudspeaker, and both talkers played through different, spatially separated loudspeakers). The results were analyzed to determine the relative performance levels that would occur with each AV configuration with target information that was equally likely to originate from either of the two talkers in the stimulus. The results indicate that spatial separation of the audio signals has the greatest impact on performance in multichannel AV speech displays and that caution should be used when presenting a visual representation of only a single talker unless that talker is known to be the highest priority talker in the combined AV stimulus. Potential applications of this research include the design of improved audiovisual speech displays for multichannel communications systems.

Adult↗

The isoazimuthal perception of sounds across distance: a preliminary investigation into the location of the audio egocenter.

Evidence indicates that both visual and auditory input may be represented in multiple frames of reference at different processing stages in the nervous system. Most models, however, have assumed that unimodal auditory input is first encoded in a head-centered reference frame. The present work tested this conjecture by measuring the subjective auditory egocenter in six blindfolded listeners who were asked to match the perceived azimuths of sounds that were alternately played between a surrounding arc of far-field speakers and a hand-held point source located three different distances from the head. If unimodal auditory representation is head centered, then "isoazimuth" lines fitted to the matching estimates across distance should intersect near the midpoint of the interaural axis. For frontomedially arranged speakers, isoazimuth lines instead converged in front of the interaural axis for all listeners, often at a point between the two eyes. As far-field sources moved outside the visual field, however, the auditory egocenter location implied by the intersection of the isoazimuth lines retreated toward or even behind the interaural axis. Physiological and behavioral evidence is used to explain this change from an eye-centered to a head-centered auditory egocenter as a function of source laterality.

Adult↗

Within-ear and across-ear interference in a dichotic cocktail party listening task: effects of masker uncertainty.

Increases in masker variability have been shown to increase the effects of informational masking in non-speech listening tasks, but relatively little is known about the influence that masker uncertainty has on the informational components of speech-on-speech masking. In this experiment, listeners were asked to extract information from a target phrase that was presented in their right ear while ignoring masking phrases that were presented in the same ear as the target phrase and in the ear opposite the target phrase. The level of masker uncertainty was varied by holding constant or "freezing" the talkers speaking the masking phrases, the semantic content used in the masking phrases, or both the talkers and the semantic content in the masking phrases within each block of 120 trials. The results showed that freezing the semantic content of the masking phrase in the target ear was the only reduction in masker uncertainty that ever resulted in a significant improvement in performance. Providing feedback after each trial improved performance overall, but did not prevent the listeners from making incorrect responses that matched the content of the frozen target-ear masking phrase. However, removing the target-ear contents corresponding to the masking phrase from the response set resulted in a dramatic improvement in performance. This suggests that the listeners were generally able to understand both of the phrases presented to the target ear, and that their incorrect responses in the task were almost entirely a result of their inability to determine which words were spoken by the target talker.

Adult↗

Informational masking caused by contralateral stimulation.

Although informational masking is thought to reflect central mechanisms, the effects are generally much stronger when the target and masker are presented to the same ear than when they are presented to different ears. However, the results of a recent study by Brungart and Simpson [J. Acoust. Soc. Am. 112, 2985-2995 (2002)] indicated that a speech masker that is presented contralateral to a speech signal can produce substantial amounts of informational masking when a second speech masker is played simultaneously in the same ear as the signal. In this study, we conducted a series of experiments that paralleled those of Brungart and Simpson but used a pure-tone signal and multitone informational maskers in a detection task. Both the signal and the maskers were played as sequences of short bursts in each observation interval. The maskers were arranged in two types of spectrotemporal patterns. One type of pattern, called "multiple-bursts same" (MBS), has previously been shown to produce very large amounts of informational masking while the other type of pattern, called "multiple-bursts different" (MBD), has been shown to produce very small amounts of informational masking. Several conditions of ipsilateral, contralateral, and combined presentation of these maskers were tested. The results showed that presentation of the MBS masker in the contralateral ear produced a substantial amount of informational masking when the MBD masker was simultaneously presented to the ipsilateral ear. The results supported the earlier findings of Brungart and Simpson indicating that listeners are unable to selectively focus their attention on a single ear in some complex dichotic listening conditions. These results suggest that this contralateral masking effect is not restricted to speech and may reflect more general limitations on processing capacity. Further, it was concluded that the magnitude of the contralateral masking effect was related both to the informational masking value of the contralateral masker and the complexity of the stimulus and/or task in the ear in which the signal was presented.

Adolescent↗

Effects of fundamental frequency and vocal-tract length changes on attention to one of two simultaneous talkers.

Three experiments used the Coordinated Response Measure task to examine the roles that differences in F0 and differences in vocal-tract length have on the ability to attend to one of two simultaneous speech signals. The first experiment asked how increases in the natural F0 difference between two sentences (originally spoken by the same talker) affected listeners' ability to attend to one of the sentences. The second experiment used differences in vocal-tract length, and the third used both F0 and vocal-tract length differences. Differences in F0 greater than 2 semitones produced systematic improvements in performance. Differences in vocal-tract length produced systematic improvements in performance when the ratio of lengths was 1.08 or greater, particularly when the shorter vocal tract belonged to the target talker. Neither of these manipulations produced improvements in performance as great as those produced by a different-sex talker. Systematic changes in both F0 and vocal-tract length that simulated an incremental shift in gender produced substantially larger improvements in performance than did differences in F0 or vocal-tract length alone. In general, shifting one of two utterances spoken by a female voice towards a male voice produces a greater improvement in performance than shifting male towards female. The increase in performance varied with the intonation patterns of individual talkers, being smallest for those talkers who showed most variability in their intonation patterns between different utterances.

Attention↗

Auditory localization in the horizontal plane with single and double hearing protection.

INTRODUCTION: Although single hearing protection devices such as earplugs or earmuffs are known to degrade sound localization, little is known about localization accuracy in double-hearing-protection conditions where both earplugs and earmuffs are worn at the same time. METHODS: Listeners wearing earplugs, earmuffs, or a combination of earplugs and earmuffs were asked to localize short (250 ms) or long (continuous) pink noise signals originating from one of 24 loudspeaker locations in the horizontal plane. RESULTS: When single hearing protection was worn, localization was reasonably accurate in the left-right dimension even when the stimuli were short in duration. When double hearing protection was worn, however, left-right localization accuracy was poor even when the stimuli were on continuously. A second experiment showed that localization accuracy with double hearing protection varied substantially across different listeners, but that it varied only slightly across refittings of the same earplugs and earmuffs on the same listener. A third experiment showed that double hearing protection impaired localization in the left-right dimension much more for narrow-band sounds at frequencies above 500 Hz than it did for narrowband sounds at frequencies at or below 250 Hz. DISCUSSION: The severe disruptions in performance that occurred when earmuffs and earplugs were worn simultaneously suggest the influence of a mechanism such as bone conduction that does not normally interfere with localization when only a single hearing protection device is used.

Adult↗

The effects of spatial separation in distance on the informational and energetic masking of a nearby speech signal.

Although many studies have shown that intelligibility improves when a speech signal and an interfering sound source are spatially separated in azimuth, little is known about the effect that spatial separation in distance has on the perception of competing sound sources near the head. In this experiment, head-related transfer functions (HRTFs) were used to process stimuli in order to simulate a target talker and a masking sound located at different distances along the listener's interaural axis. One of the signals was always presented at a distance of 1 m, and the other signal was presented 1 m, 25 cm, or 12 cm from the center of the listener's head. The results show that distance separation has very different effects on speech segregation for different types of maskers. When speech-shaped noise was used as the masker, most of the intelligibility advantages of spatial separation could be accounted for by spectral differences in the target and masking signals at the ear with the higher signal-to-noise ratio (SNR). When a same-sex talker was used as the masker, the intelligibility advantages of spatial separation in distance were dominated by binaural effects that produced the same performance improvements as a 4-5-dB increase in the SNR of a diotic stimulus. These results suggest that distance-dependent changes in the interaural difference cues of nearby sources play a much larger role in the reduction of the informational masking produced by an interfering speech signal than in the reduction of the energetic masking produced by an interfering noise source.

Adult↗

Within-ear and across-ear interference in a cocktail-party listening task.

Although many researchers have shown that listeners are able to selectively attend to a target speech signal when a masking talker is present in the same ear as the target speech or when a masking talker is present in a different ear than the target speech, little is known about selective auditory attention in tasks with a target talker in one ear and independent masking talkers in both ears at the same time. In this series of experiments, listeners were asked to respond to a target speech signal spoken by one of two competing talkers in their right (target) ear while ignoring a simultaneous masking sound in their left (unattended) ear. When the masking sound in the unattended ear was noise, listeners were able to segregate the competing talkers in the target ear nearly as well as they could with no sound in the unattended ear. When the masking sound in the unattended ear was speech, however, speech segregation in the target ear was substantially worse than with no sound in the unattended ear. When the masking sound in the unattended ear was time-reversed speech, speech segregation was degraded only when the target speech was presented at a lower level than the masking speech in the target ear. These results show that within-ear and across-ear speech segregation are closely related processes that cannot be performed simultaneously when the interfering sound in the unattended ear is qualitatively similar to speech.

Adult↗