PubMed HealthSearch

Biomedical subjects

F G Zeng

Publications and source records attributed to F G Zeng.

At least 19 recordsLinked to original sources

Amplitude mapping and phoneme recognition in cochlear implant listeners.

OBJECTIVE: Speech and other environmental sounds must be compressed to accommodate the small electric dynamic range in cochlear implant listeners. The objective of this paper is to study whether and how amplitude compression and dynamic range reduction affect phoneme recognition in quiet and in noise for cochlear implant listeners. DESIGN: Four implant listeners using the Nucleus-22 SPEAK speech processor participated in this study. The amount of compression was varied by manipulating the Q-value in the SPEAK processor. The size of the dynamic range was systematically reduced by increasing the threshold level and decreasing the comfortable level in the processor. Both female- and male-talker vowel and consonant materials were used to evaluate speech recognition performance in quiet and in noise. Speech-spectrum-shaped noise was mixed with the speech signal and presented continuously to the speech processor through a direct electric connection. Signal to noise ratios were changed over a 30 to 40 dB range, within which phoneme recognition increased from chance to asymptotic performance. Phoneme recognition scores were obtained as the number of active electrodes was reduced from 20 to 10 to 4. For purposes of comparison, phoneme recognition data also were collected in four normal-hearing listeners under comparable laboratory conditions. RESULTS: In both quiet and noise, the amount of amplitude compression did not significantly affect phoneme recognition. The reduction of dynamic range marginally affected phoneme recognition in quiet, but significantly degraded phoneme recognition in noise. Generally, the 20- and 10-electrode processors produced similar performance, whereas the 4-electrode processor produced significantly poorer performance. Compared with normal-hearing listeners, cochlear-implant listeners required higher signal to noise ratios to achieve comparable recognition performance and produced significantly lower recognition scores at the same signal to noise ratios. CONCLUSIONS: The amount of amplitude compression does not significantly affect phoneme recognition, whereas reducing dynamic range significantly lowers phoneme recognition, particularly in noise and for vowels. Because the SPEAK processor extracts mostly spectral peaks, the present conclusions may not be applied to other types of processors extracting temporal envelope cues. The present results also suggest that more than four electrodes are required to optimize speech recognition in multiple-talker and noise conditions. A significant performance gap in speech recognition still remains between cochlear implant and normal-hearing listeners at the same signal to noise ratios. Improved cochlear implant designs and fitting procedures are required to narrow and, hopefully, close this performance gap.

Adult

Encoding loudness by electric stimulation of the auditory nerve.

Electric charge has long been hypothesized to be the effective stimulus variable that determines loudness evoked by directly stimulating the auditory nerve. This 'equal-charge, equal-loudness' hypothesis predicts that stimulus amplitude and duration can be traded linearly to produce equal loudness. Loudness sensations from threshold to maximum loudness were measured systematically as a function of stimulus amplitude and duration in cochlear implant listeners. The measured data do not support the equal-charge, equal-loudness hypothesis: an increment in stimulus amplitude produces a significantly louder sensation than the same change in stimulus duration. Instead of the linear equal-charge model, a power-function model successfully predicts the measured data and should be used to encode loudness in electric hearing.

Adult

Design for an inexpensive but effective cochlear implant.

Widespread application of cochlear implants is limited by cost, especially in developing countries. In this article we present a design for a low-cost but effective cochlear implant system. The system includes a speech processor, four pairs of transmitting and receiving coils, and an electrode array with four monopolar electrodes. All implanted components are passive, reducing to a minimum the complexity of manufacture and allowing high reliability. A four-channel continuous interleaved sampling strategy is used for the speech processor. The processor and transmission link have been evaluated in tests with a subject previously implanted with the Ineraid electrode array and percutaneous connector. A prototype of the link, consisting of four pairs of transmitting and external receiving coils, was used, with the outputs of the receiving coils directed to four intracochlear electrodes through the percutaneous connector. The subject achieved speech reception scores with the prototype system that were equivalent to those achieved with a standard laboratory implementation of a continuous interleaved sampling processor with current-controlled stimuli.

Cochlear Implantation

Interactions of forward and simultaneous masking in intensity discrimination.

Intensity coding mechanisms are explored in a paradigm involving both forward and simultaneous masking. For intensity discrimination of 1000-Hz pure tone in quiet, a near-miss to Weber's law is observed. However, as more stimulus components are added to this relatively simple experiment, interactions among components produce a more complex pattern of results. An intense forward masker, while not causing any threshold shift for the test tone, produces a nonmonotonic intensity discrimination function ["the midlevel hump," Zeng et al., Hearing Res. 55, 223-230 (1991)]. The midlevel hump can be removed by the presence of additional notched noise [Plack and Viemeister, J. Acoust. Soc. Am. 92, 1902-1910 (1992)] or narrow-band noise whose level is increased along with the test tone's standard level. The same midlevel hump can also be enhanced by a fixed-low-level notched noise or a high-level, high-pass noise which causes minimal masking at the test frequency. Interactions of forward masking and simultaneous masking present a serious problem for a clear interpretation of these results. For example, the notched noise was originally intended to restrict off-frequency listening, but on-frequency masking compromised this original purpose and confounded the interpretation of the notched noise effects. By measuring systematically the growth-of-masking functions, the present study identified various interactions of forward and simultaneous masking and clarified the role of off-frequency listening in forward-masked intensity discrimination. Both peripheral and central mechanisms may have contributed to the occurrence, reduction and enhancement of the midlevel hump under these masking conditions.

Adult

Importance of tonal envelope cues in Chinese speech recognition.

Recent studies have shown that temporal waveform envelope cues can provide significant information for English speech recognition. This study investigated the use of temporal envelope cues in a tonal language: Mandarin Chinese. In this study, the speech was divided into several frequency analysis bands; the amplitude envelope was extracted from each band by half-wave rectification and low-pass filtering and was used to modulate a noise of the same bandwidth as the analysis band. These manipulations preserved temporal and amplitude cues in each frequency band, but removed the spectral detail within each band. Chinese vowels, consonants, tones and sentences were identified by 12 native Chinese-speaking listeners with 1, 2, 3, and 4 noise bands. The results showed that the recognition score of vowels, consonants, and sentences increased monotonically with the number of bands, a pattern similar to that observed in English speech recognition. In contrast, tones were consistently recognized at about 80% correct level, independent of the number of bands. This high level of tone recognition produced a significant difference in the open-set sentence recognition between Chinese (11.0%) and English (2.9%) for the one-band condition where no spectral information was available. The data also revealed that, with primarily temporal cues, the falling-rising tone (tone 3) and the falling tone (tone 4) were more easily recognized than the flat tone (tone 1) and the rising tone (tone 2). This differential pattern in tone recognition resulted in a similar pattern in word recognition: words having either tone 3 or 4 were more likely to be recognized while words having tone 1 and 2 were not. The quantitative role of tones in Chinese speech recognition was further explored using a power-function model and found to play a significant role in relating phoneme recognition to sentence recognition.

Adult

Loudness of dynamic stimuli in acoustic and electric hearing.

Traditional loudness models have been based on the average energy and the critical band analysis of steady-state sounds. However, most environmental sounds, including speech, are dynamic stimuli, in which the average level [e.g., the root-mean-square (rms) level] does not account for the large temporal fluctuations. The question addressed here was whether two stimuli of the same rms level but different peak levels would produce an equal loudness sensation. A modern adaptive procedure was used to replicate two classic experiments demonstrating that the sensation of "beats" in a two- or three-tone complex resulted in a louder sensation [E. Zwicker and H. Fastl, Psychoacoustics-Facts and Models (Springer-Verlag, Berlin, 1990)]. Two additional experiments were conducted to study exclusively the effects of the temporal envelope on the loudness sensation of dynamic stimuli. Loudness balance was performed by normal-hearing listeners between a white noise and a sinusoidally amplitude-modulated noise in one experiment, and by cochlear implant listeners between two harmonic stimuli of the same magnitude spectra, but different phase spectra, in the other experiment. The results from both experiments showed that, for two stimuli of the same rms level, the stimulus with greater temporal fluctuations sometimes produced a significantly louder sensation, depending on the temporal frequency and overall stimulus level. In normal-hearing listeners, the louder sensation was produced for the amplitude-modulated stimuli with modulation frequencies lower than 400 Hz, and gradually disappeared above 400 Hz, resulting in a low-pass filtering characteristic which bore some similarity to the temporal modulation transfer function. The extent to which loudness was greater was a nonmonotonic function of level in acoustic hearing and a monotonically increasingly function in electric hearing. These results suggest that the loudness sensation of a dynamic stimulus is not limited to a 100-ms temporal integration process, and may be determined jointly by a compression process in the cochlea and an expansion process in the brain. A level-dependent compression scheme that may better restore normal loudness of dynamic stimuli in hearing aids and cochlear implants is proposed.

Acoustics

Distortion product otoacoustic emission suppression tuning curves in human adults and neonates.

Distortion product otoacoustic emission (DPOAE) iso-suppression tuning curves (STC) were generated in 15 normal-hearing adults and 16 healthy term-born neonates for three f2 frequencies. The 2f1-f2 DPOAE was elicited using f2/f1 = 1.2, LI = 1.2, LI = 65 and L2 = 50 dB SPL. A suppressor tone was presented at frequencies ranging from 1 octave below to 1/4 octave above f2 and varied in level until DPOAE amplitude was reduced by 6 dB. The suppressor level required for 6 dB suppression was plotted as function of suppressor frequency to generate a DPOAE STC. Forward-masked psychoacoustic tuning curves (PTC) were obtained for three of the adult subjects. Results indicate that DPOAE STCs are stable and show minimal inter- and intra-subject variability. The tip of the STC is consistently centered around the f2 region and STCs are similar in shape, width (Q10) and slope to VIIIth-nerve TCs. PTCs and STCs measured in the same subject showed similar trends, although PTCs had narrower width and steeper slope. Neonatal STCs were recorded at 3000 and 6000 Hz only and were comparable in shape, width and slope to adult STCs. Results suggest: (1) suppression of the 2f1-f2 DPOAE may provide an indirect measure of cochlear frequency resolution in humans and (2) cochlear tuning, and associated active processes in the cochlea, are mature by term birth for at least mid- and high-frequencies. These results provide significant impetus for continued study of DPOAE suppression as a means of evaluating cochlear frequency resolution in humans.

Acoustic Stimulation

Speech recognition with primarily temporal cues.

Nearly perfect speech recognition was observed under conditions of greatly reduced spectral information. Temporal envelopes of speech were extracted from broad frequency bands and were used to modulate noises of the same bandwidths. This manipulation preserved temporal envelope cues in each band but restricted the listener to severely degraded information on the distribution of spectral energy. The identification of consonants, vowels, and words in simple sentences improved markedly as the number of bands increased; high speech recognition performance was obtained with only three bands of modulated noise. Thus, the presentation of a dynamic temporal pattern in only a few broad spectral regions is sufficient for the recognition of speech.

Auditory Threshold

Possible origins of the non-monotonic intensity discrimination function in forward masking.

A non-monotonic intensity discrimination function in forward masking has been recently reported [Zeng et al. (1991) Hear. Res. 55, 223-230; Zeng and Turner (1992) J. Acoust Soc. Am. 92, 782-787] in which just-noticeable-differences (jnds) in intensity are largest for midlevel tones and smaller for soft and loud tones following an intense narrow-band noise. One hypothesis was that this midlevel hump reflects the contribution of low-spontaneous rate (SR) neurons to intensity coding, based on the differential recovery from forward masking of low-SR and high-SR neurons [Relkin and Doucet (1991) Hear. Res. 55, 215-222]. The present study conducted three experiments stimulating different stages of the auditory system in an attempt to determine the peripheral and central origins of the midlevel hump. First, in two cochlear implant (CI) listeners, the forward masker produced a midlevel hump on the intensity discrimination function, suggesting that the synapses between the hair cell and the eighth nerve are probably not responsible for the hump, as they are bypassed and the eighth nerve is stimulated directly. Second, in auditory brainstem implant (ABI) listeners, the forward masker produced no midlevel hump, but the masked jnds were larger than those without a masker. The absence of the midlevel hump in the ABI listeners suggests that the occurrence of the hump requires physiological mechanisms in the auditory nerve transmission, or the intrinsic processing circuits of the cochlear nuclei, or both. Third, in normal-hearing listeners, an ipsilateral, 90 dB SPL, pure-tone forward masker produced a midlevel hump, which is similar to that using a narrow-band noise masker; whereas a contralateral forward masker produced essentially no midlevel hump, suggesting that binaural interactions at superior olivary complex and more central sites are probably not responsible.

Acoustic Stimulation

Loudness of simple and complex stimuli in electric hearing.

In a earlier study we showed that loudness function depends on stimulus frequency for simple sinusoidal and pulsatile stimuli in cochlear implants. Loudness is an exponential function of stimulus amplitude for high frequencies (> 300 Hz) and a power function for low frequencies (< 300 Hz). Two experiments were conducted to extend our previous work in eight Ineraid cochlear implant subjects. First, loudness functions were measured by means of a magnitude estimation technique for simple sinusoidal stimuli. Stimuli were 100- and 1,000-Hz sinusoids. Stimuli had a duration of 300 milliseconds and were presented at amplitudes representing 10%, 30%, 50%, 70%, and 90% of the dynamic range. Loudness estimates were best fit by an exponential function for the 1,000-Hz sinusoid and by a power function for the 100-Hz sinusoid. The estimation result is consistent with previous results that were obtained with a loudness balance technique. In a second experiment, loudness functions were studied for modulated stimuli, in which a 1,000-Hz sinusoid was modulated by a 100-Hz sinusoid at either 50% or 100% modulation depth. A linear loudness-balance function was obtained between the modulated stimuli and a 1,000-Hz sinusoidal standard, indicating that the modulated stimuli have the same exponential loudness function as the 1,000-Hz carrier despite the 100-Hz modulator. This result has important implications for speech processor design, because many devices use low-frequency speech envelopes to modulate a high-frequency carrier. This result also provides a psychophysical basis for the logarithmic amplitude transformation in commonly used speech processors to compress the wide acoustic dynamic range into a narrow electric dynamic range while preserving normal loudness function.

Adult

Loudness-coding mechanisms inferred from electric stimulation of the human auditory system.

Two distinct physiological mechanisms underlying loudness sensation were inferred from electric stimulation of the human auditory nerve and brainstem. In contrast to a power function relating loudness and stimulus intensity in acoustic hearing, loudness in electric stimulation of the auditory nerve depends on stimulus frequency. Loudness is an exponential function of electric amplitude for high frequencies and is a power function for low frequencies. A frequency-dependent, two-stage model is suggested to explain the loudness function, in which the first stage of processing is performed by a mechanical mechanism in the cochlea for high-frequency stimuli and by a neural mechanism in the cochlear nucleus for low-frequency stimuli.

Acoustic Stimulation

Loudness growth in forward masking: relation to intensity discrimination.

The growth of loudness of a tone burst following an intense forward masker was measured as a function of the tone level. The level of the forward-masked tone was adjusted to balance the loudness of a standard tone presented without a forward masker, using a 2AFC, double-staircase, tracking procedure. The forward masker was a 90-dB SPL, 100-ms, 1000-Hz pure tone. The standard tone and the masked tone were both 25-ms, 1000-Hz pure tones. The forward masker and the masked tone were always presented in the first interval. With a 100-ms delay between them, there was little or no threshold elevation for the masked tone. However, the masker caused the masked tone to sound louder than it would if it had not been masked, a phenomenon termed "loudness enhancement" [Irwin and Zwislocki, Percept. Psychophys. 10, 189-192 (1971); Galambos et al., J. Acoust. Soc. Am. 52, 1127-1130 (1972)]. In addition, the present results show a nonmonotonic enhancement function that the forward masker introduced a 10-16-dB enhancement effect for tones of 40-65 dB SPL and no significant effect for the 30 and 90 dB SPL tones. The loudness variability in forward masking was estimated from the upper and lower sequences tracking the 21% and the 79% louder response levels on the psychometric function, respectively. The variability demonstrated a similar nonmonotonic function. In forward masking loudness grows more steeply at low-medium sensation levels, and merges with normal growth at high levels.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult

Loudness balance between electric and acoustic stimulation.

Binaural loudness balance between electric and acoustic stimulation is obtained in auditory brainstem implant listeners who had substantial acoustic hearing in one ear. The data are well described by a linear relationship between acoustic decibels and electric microamps. Based upon this linear relationship, we propose an exponential model of loudness growth in electric stimulation. The exponential model predicts that the loudness growth function can be determined solely by the threshold and the uncomfortable loudness level in electric stimulation. This prediction is consistent with previous psychophysical data on loudness functions. Implications of this finding for speech processor designs are discussed.

Acoustic Stimulation

Intensity discrimination in forward masking.

A nonmonotonic intensity discrimination function was recently reported in which a midlevel hump occurred for 25-ms sinusoidal standards ranging from 20 to 100 dB SPL and presented 100 ms after an intense narrow-band noise forward masker [F.-G. Zeng et al., Hear. Res. 55, 223-230 (1991)]. This paper provides additional data on how the midlevel hump is affected by three factors of forward masking: signal delay, masker level, and frequency. Specifically, just-noticeable differences (jnd's) in intensity were obtained at signal delays of 50, 200, and 400 ms. Results show that at the midlevels the forward-masked intensity jnd's did not recover to their unmasked values, even at the 400-ms signal delay. The longer the signal delay, the smaller this midlevel hump. This slow recovery of the midlevel jnd's is consistent with the finding that low-spontaneous rate (SR) neurons have a slow recovery from forward masking [E. M. Relkin and J. R. Doucet, Hear. Res. 55, 215-222 (1991)]. The large midlevel effect decreased sharply as masker level was reduced from 90 to 60 dB SPL, and disappeared for masker levels less than 40 dB SPL. A frequency selectivity effect for the large midlevel jnd effect was also observed, as maskers with frequency components 2 to 3 oct away from the signal frequency did not affect the jnd's. Overall, the present data are consistent with the hypothesis of Zeng et al. (1991) that low-SR neurons are involved in the midlevel hump of intensity discrimination in forward masking.

Adult

Frequency discrimination in forward and backward masking.

Frequency difference limens for pure tones preceded by a forward masker or followed by a backward masker were obtained across a wide range of signal levels. Relkin and Doucet [Hear. Res. 55, 215-222 (1991)] have shown that at a masker-signal delay of 100 ms, the thresholds of high-SR (spontaneous rate) auditory-nerve fibers are recovered, while the low-SR fiber thresholds are not. Therefore, forward-masked frequency discrimination potentially offers a method to investigate the role of low-SR fibers in the coding of frequency. It has been shown that when an intense forward masker is presented 100 ms before a pure-tone signal, intensity difference limens are elevated for mid-level signals [Zeng et al., Hear. Res. 55, 223-230 (1991)]. However, Plack and Viemeister [J. Acoust. Soc. Am. 92, 3097-3101 (1992)] have shown that a similar elevation in the intensity difference limen is obtained under conditions of backward masking, where selective adaptation of the auditory neurons would not be expected to occur. A condition of backward-masked frequency discrimination was therefore included to investigate the role of interference resulting from adding additional stimuli to a discrimination task. For signals at 1000 and 6000 Hz, there was no effect of a forward masker upon frequency difference limens. For the backward-masked conditions, an elevation of the frequency difference limen was observed at all signal levels, demonstrating that the effects of forward and backward maskers upon frequency discrimination are dissimilar and suggesting that cognitive effects are present in backward-masked discrimination tasks.(ABSTRACT TRUNCATED AT 250 WORDS)

Auditory Perception

Recovery from prior stimulation. II: Effects upon intensity discrimination.

We obtained just-noticeable differences (jnds) for the intensity of pure tones following a forward masker. The masker was a 100-ms burst of narrow-band noise centered at 1000 Hz presented at 90 dB SPL; the pure-tone signal was at 1000 Hz and was 25 ms in duration. The masker-signal delay was 100 ms. Under these conditions, there is no threshold shift for the detection of the pure-tone signal following the forward masker. In contrast with the absence of a forward-masker effect upon detection thresholds, unusually large midlevel (40-60 dB SPL) jnds were observed. These large midlevel jnds were measured as a function of signal delay, revealing that they are not completely recovered to the normal (unmasked) values by 400 ms. We interpret these data as a consequence of the slower recovery of low-spontaneous rate, high-threshold neurons following prior stimulation (Relkin and Doucet, 1990). These experiments may therefore provide psychophysical evidence that the low-spontaneous rate, high-threshold neurons are a necessary physiological component in the coding of the large dynamic range for intensity. In addition, the present data provide evidence that the assumption that the effect of forward masking is limited to 100-200 ms is inappropriate, as this recovery time does not necessarily apply to suprathreshold tasks.

Acoustic Stimulation

Binaural loudness matches in unilaterally impaired listeners.

Binaural loudness matching data using a 21FC adaptive procedure were obtained in high-frequency, unilateral cochlear-impaired listeners. The matches were obtained at frequencies where both ears had similarly normal thresholds, and also at other frequencies where the impaired ear had various degrees of hearing loss. In these listeners, one presumed difference between the ears is the limited or altered spread of excitation in the impaired ear. In agreement with previous studies using other approaches (Hellman, 1974, 1978; Hellman & Meiselman, 1986; Moore, Glasberg, Hess & Birchall, 1985; Schneider & Parker, 1987), the results of the present study suggest that both the range and the slope of loudness growth function are not dependent on the spread of excitation, but instead are related primarily to the degree of threshold elevation at the test frequency. Following this suggestion, a spread-of-excitation-independent model, based upon a group of neurons with the same characteristic frequency (CF) but different thresholds, is proposed to account for loudness growth in both normal and recruitment cases. In particular, it is shown quantitatively that a compressed distribution of thresholds due to threshold elevation may be responsible for loudness recruitment in sensorineural hearing loss.

Adult

Recognition of voiceless fricatives by normal and hearing-impaired subjects.

The purpose of this study was to investigate the sufficient perceptual cues used in the recognition of four voiceless fricative consonants [s, f, theta, integral of] followed by the same vowel [i:] in normal-hearing and hearing-impaired adult listeners. Subjects identified the four CV speech tokens in a closed-set response task across a range of presentation levels. Fricative syllables were either produced by a human speaker in the natural stimulus set, or generated by a computer program in the synthetic stimulus set. By comparing conditions in which the subjects were presented with equivalent degrees of audibility for individual fricatives, it was possible to isolate the factor of lack of audibility from that of loss of suprathreshold discriminability. Results indicate that (a) the friction burst portion may serve as a sufficient cue for correct recognition of voiceless fricatives by normal-hearing subjects, whereas the more intense CV transition portion, though it may not be necessary, can also assist these subjects to distinguish place information, particularly at low presentation levels; (b) hearing-impaired subjects achieved close-to-normal recognition performance when given equivalent degrees of audibility of the frication cue, but they obtained poorer-than-normal performance if only given equivalent degrees of audibility of the transition cue; (c) the difficulty that hearing-impaired subjects have in perceiving fricatives under normal circumstances may be due to two factors: the lack of audibility of the frication cue and the loss of discriminability of the transition cue.

Adult