Auditive and cognitive factors in speech perception by elderly listeners. III. Additional data and final discussion.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to R Plomp.
Explore the source record for details and available documents.
The effect of reduced spectral contrast on the speech-reception threshold (SRT) for sentences in noise and on phoneme identification, was investigated with 16 normal-hearing subjects. Signal processing was performed by smoothing the envelope of the squared short-time fast Fourier transform (FFT) by convolving it with a Gaussian-shaped filter, and overlapping additions to reconstruct a continuous signal. Spectral energy in the frequency region from 100 to 8000 Hz was smeared over bandwidths of 1/8, 1/4, 1/3, 1/2, 1, 2, and 4 oct for the SRT experiment. Vowel and consonant identification was studied for smearing bandwidths of 1/8, 1/2, and 2 oct. Results showed the SRT in noise to increase as the spectral energy was smeared over bandwidths exceeding the ear's critical bandwidth. Vowel identification suffered more from this type of processing than consonant identification. Vowels were primarily confused with the back vowels /c,u/, and consonants were confused where place of articulation is concerned.
Speech-reception thresholds (SRT) were measured for 17 normal-hearing and 17 hearing-impaired listeners in conditions simulating free-field situations with between one and six interfering talkers. The stimuli, speech and noise with identical long-term average spectra, were recorded with a KEMAR manikin in an anechoic room and presented to the subjects through headphones. The noise was modulated using the envelope fluctuations of the speech. Several conditions were simulated with the speaker always in front of the listener and the maskers either also in front, or positioned in a symmetrical or asymmetrical configuration around the listener. Results show that the hearing impaired have significantly poorer performance than the normal hearing in all conditions. The mean SRT differences between the groups range from 4.2-10 dB. It appears that the modulations in the masker act as an important cue for the normal-hearing listeners, who experience up to 5-dB release from masking, while being hardly beneficial for the hearing impaired listeners. The gain occurring when maskers are moved from the frontal position to positions around the listener varies from 1.5 to 8 dB for the normal hearing, and from 1 to 6.5 dB for the hearing impaired. It depends strongly on the number of maskers and their positions, but less on hearing impairment. The difference between the SRTs for binaural and best-ear listening (the "cocktail party effect") is approximately 3 dB in all conditions for both the normal-hearing and the hearing-impaired listeners.
The rationale for a method to quantify the information content of linguistic stimuli, i.e., the linguistic entropy, is developed. The method is an adapted version of the letter-guessing procedure originally devised by Shannon [Bell Syst. Tech. J. 30, 50-64 (1951)]. It is applied to sentences included in a widely used test to measure speech-reception thresholds and originally selected to be approximately equally redundant. Results of a first experiment reveal that this method enables one to detect subtle differences between sentences and sentence lists with respect to linguistic entropy. Results of a second experiment show that (1) in young listeners and with the sentences employed, manipulating linguistic entropy can result in an effect on SRT of approximately 4 dB in terms of signal-to-noise ratio; (2) the range of this effect is approximately the same in elderly listeners.
Within a study on the merits of a multichannel automatic gain control in hearing aids, the effect of frequency-selective amplification on the masked speech-reception threshold (SRT) for sentences is measured in conditions of seriously disturbing low-frequency noise, with the effect of wideband amplification as a reference. Speech and noise are both spectrally shaped according to the bisector line of the listener's dynamic-range of hearing, but with the noise in a single octave band (0.25-0.5 or 0.5-1 kHz) increased by 20 dB relative to this line. The increase of noise level is steady state in the first experiment, and time varying in the second experiment. Results for 12 normal-hearing and 12 hearing-impaired listeners indicate that, in both experiments, frequency-selective compression of the signal in the octave band with the 20-dB increase of noise is more beneficial than wideband compression. For the hearing-impaired group, wideband compression does not give any systematic change in intelligibility. Frequency-selective compression in steady-state conditions may, for both groups of listeners, give a decrease of masked SRT (relative to a condition without compression) of up to 4 dB for a compression factor of 100%. Roughly comparable effects are seen for frequency-selective compression in time-varying conditions. The superiority of frequency-selective over wideband compression is attributed to a more effective reduction of upward spread of masking.
In part I of this study [van Rooij et al., J. Acoust. Soc. Am. 86, 1294-1309 (1989)], the validity and manageability of a test battery comprising auditive (sensitivity, frequency resolution, and temporal resolution), cognitive (memory performance, processing speed, and intellectual abilities), and speech perception tests (at the phoneme, spondee, and sentence level) were investigated. In the present article, the results of a selection of these tests for 72 elderly subjects (aged 60-93 years) are analyzed by multivariate statistical techniques. The results show that the deterioration of speech perception in the elderly consists of two statistically independent components: (a) a large component mainly representing the progressive high-frequency hearing loss with age that accounts for approximately two-thirds of the systematic variance of the tests of speech perception and (b) a smaller component (accounting for one-third of the systematic variance of the speech perception tests) mainly representing a general performance decrement due to reduced mental efficiency, which is indicated by a general slowing of performance and a reduced memory capacity. Although both components are correlated with age, it was found that the balance between auditive and cognitive contributions to speech perception performance did not change with age.
The speech-reception threshold (SRT) for sentences presented in a fluctuating interfering background sound of 80 dBA SPL is measured for 20 normal-hearing listeners and 20 listeners with sensorineural hearing impairment. The interfering sounds range from steady-state noise, via modulated noise, to a single competing voice. Two voices are used, one male and one female, and the spectrum of the masker is shaped according to these voices. For both voices, the SRT is measured as well in noise spectrally shaped according to the target voice as shaped according to the other voice. The results show that, for normal-hearing listeners, the SRT for sentences in modulated noise is 4-6 dB lower than for steady-state noise; for sentences masked by a competing voice, this difference is 6-8 dB. For listeners with moderate sensorineural hearing loss, elevated thresholds are obtained without an appreciable effect of masker fluctuations. The implications of these results for estimating a hearing handicap in everyday conditions are discussed. By using the articulation index (AI), it is shown that hearing-impaired individuals perform poorer than suggested by the loss of audibility for some parts of the speech signal. Finally, three mechanisms are discussed that contribute to the absence of unmasking by masker fluctuations in hearing-impaired listeners. The low sensation level at which the impaired listeners receive the masker seems a major determinant. The second and third factors are: reduced temporal resolution and a reduction in comodulation masking release, respectively.
A key issue in research on speech perception in the elderly is whether the difficulties in understanding speech are caused by auditive and/or cognitive factors. Resolving this issue is not only scientific interest but has many practical, i.e. diagnostic and rehabilitative, implications as well. We developed a test battery comprising auditive (sensitivity, frequency selectivity and temporal resolution), cognitive (memory performance, processing speed and intellectual abilities), and speech perception tests (at the phoneme, spondee and sentence level). This test battery was administered to 72 elderly subjects (aged 60 to 93 years). The results show that the deterioration of speech perception in the elderly consists of two statistically independent components (a) a large component mainly representing the progressive high-frequency hearing loss with age which accounts for approximately two-thirds of the systematic variance of the tests of speech perception, and (b) a smaller component (accounting for one-third of the systematic variance of the speech perception tests) mainly representing a general performance decrement due to reduced mental efficiency, which is indicated by a general slowing of performance and a reduced memory capacity. Although both components are correlated with age, it was found that the balance between auditive and cognitive contributions to speech perception performance did not change with age.
The present paper describes a clinical test for the assessment of speech perception in noise. The test was designed to separate the effects of several relevant monaural and binaural cues. Results show that the performance of individual hearing-impaired listeners deviates significantly from normal for at least 2 of the following aspects: (1) perception of speech in steady-state noise; (2) relative binaural advantage due to directional cues; (3) relative advantage due to masker fluctuations. In contrast, both the hearing loss for reverberated speech and the relative binaural advantage due to interaural signal decorrelation, caused by reverberation, were essentially normal for almost all hearing impaired.
A group of 15 patients with complaints of having difficulties in understanding speech, especially in noisy surroundings in spite of (nearly) normal pure-tone audiograms, was subjected to a battery of speech-audiometric tests. The results showed that these subjects had a statistically significantly higher speech reception threshold (SRT) for sentences in noise than a reference group of 10 normal-hearing subjects. This difference was most clear for a fluctuating masking noise. In conditions with much reverberation, the patients also proved to be handicapped more than the control group. Binaural hearing gain was equal for both groups. The pathogenesis of the speech-hearing loss is not known, but assessment of the SRT in noise proves to be a valuable asset in objectifying these patients' complaints.
Most sensorineurally hearing-impaired listeners need a better signal-to-noise ratio for speech reception than normal-hearing listeners do. This aspect of hearing loss is probably related to deterioration of the signal-analysing power of the ear. In an impaired ear, noise in one frequency region may have considerable masking effects in other frequency regions. To prevent the hearing-impaired listener from excessive masking, we need an adaptive hearing aid that selectively amplifies only those frequency bands with a signal quality that can contribute to intelligibility. It suffices to present signals in other frequency regions at a just-audible level. A signal-processing scheme is proposed that meets these requirements. The audiological reasoning for the spectral characteristics of such a hearing aid is illustrated with the Articulation Theory. The temporal characteristics required for adaptation to changes in the acoustic environment are discussed in terms of the Modulation Transfer Function. It is shown that amplitude compression with short time constants seriously reduces the quality of speech transmission as expressed in the Speech Transmission Index.
In a study on the effects of a frequency-dependent automatic gain control in hearing-aids, two experiments were carried out with hearing-impaired listeners. In the first experiment, the effect of varying the amplitude-frequency response on the speech-reception threshold (SRT) of sentences presented in noise was studied. The noise had the same spectrum as the long-term average spectrum of the sentences. Results suggest that the amplitude-frequency response may vary within a range from, roughly, -3 to +10 dB/oct relative to the bisector of the dynamic range, without giving an increase in SRT larger than 2 dB. In the second experiment, the effect on masked SRT of adjusting the amplitude-frequency response to situations of seriously interfering low-frequency noise was studied. Again, the noise had a spectrum identical with the long-term average spectrum of the sentences, but this time the noise level in one octave band was increased by 20 dB. Preliminary results indicate that a selective attenuation of the signal in the band containing the extra noise may give a decrease of masked SRT up to 4.5 dB.
In an evaluation of frequency-dependent automatic gain-control systems in hearing aids, the effect of varying the amplitude-frequency response on the speech-reception threshold (SRT) for sentences in noise is studied for 20 hearing-impaired listeners. The noise has a spectrum identical to the long-term average spectrum of the sentences. Speech and noise are shaped by the same amplitude-frequency response; their spectra are varied relative to the bisector of the individual's dynamic range. In four experimental conditions, the effect of a steady-state amplitude-frequency response is studied. Steepening the negative spectral slope of speech and noise appears to cause an increase of masked SRT, possibly due to increased effect of upward spread of masking. The effect of a single transition of the amplitude-frequency response between 10 and -10 dB/oct halfway through the sentence seems to be related to the effect for the fixed -10-dB/oct condition. Two transition times are tested. For a transition time of 0.25 s, the SRT is only a little higher than for 1 s. The results suggest that the amplitude-frequency response may be varied in time without having a detrimental effect on the masked SRT of sentences for hearing-impaired listeners as long as strongly negatively sloping spectra are avoided.
The effect of head-induced interaural time delay (ITD) and interaural level differences (ILD) on binaural speech intelligibility in noise was studied for listeners with symmetrical and asymmetrical sensorineural hearing losses. The material, recorded with a KEMAR manikin in an anechoic room, consisted of speech, presented from the front (0 degree), and noise, presented at azimuths of 0 degree, 30 degrees, and 90 degrees. Derived noise signals, containing either only ITD or only ILD, were generated using a computer. For both groups of subjects, speech-reception thresholds (SRT) for sentences in noise were determined as a function of: (1) noise azimuth, (2) binaural cue, and (3) an interaural difference in overall presentation level, simulating the effect of a monaural hearing acid. Comparison of the mean results with corresponding data obtained previously from normal-hearing listeners shows that the hearing impaired have a 2.5 dB higher SRT in noise when both speech and noise are presented from the front, and 2.6-5.1 dB less binaural gain when the noise azimuth is changed from 0 degree to 90 degrees. The gain due to ILD varies among the hearing-impaired listeners between 0 dB and normal values of 7 dB or more. It depends on the high-frequency hearing loss at the side presented with the most favorable signal-to-noise (S/N) ratio. The gain due to ITD is nearly normal for the symmetrically impaired (4.2 dB, compared with 4.7 dB for the normal hearing), but only 2.5 dB in the case of asymmetrical impairment. When ITD is introduced in noise already containing ILD, the resulting gain is 2-2.5 dB for all groups. The only marked effect of the interaural difference in overall presentation level is a reduction of the gain due to ILD when the level at the ear with the better S/N ratio is decreased. This implies that an optimal monaural hearing aid (with a moderate gain) will hardly interfere with unmasking through ITD, while it may increase the gain due to ILD by preventing or diminishing threshold effects.
This study compares performance of 24 young normal-hearing (aged 18-28 years) and 24 elderly (aged 61-85 years) listeners on auditive (sensitivity, frequency selectivity, and temporal resolution), cognitive (memory performance, processing speed, and divided attention ability), and speech perception tests (at the phoneme, spondee, and sentence level). Its principal aim is to assess whether the tests selected yield meaningful results. The results obtained will be used to reduce the test battery in order to be manageable in a second study on a much larger number of elderly listeners. The relationships between the tests are explored by multivariate statistical methods. The results show that: (a) in young listeners, individual differences in speech perception performance are remarkably small resulting in low correlations between the tests, while in the elderly tests of phoneme, spondee, and sentence perception overlap considerably; (b) speech perception in the elderly seems to be largely determined by hearing loss at the higher frequencies, whereas the effects of other auditive and cognitive factors seem to be relatively small or absent; and (c) performance in the elderly is only partly correlated with age.
Speech intelligibility in noise was tested pre- and postoperatively after 27 stapedectomy cases and 10 cases of reexploration after previous stapes surgery. We did not find a postoperative threshold shift (signal-to-noise ratio) for the intelligibility of sentences presented in noise. This finding corresponds well with the absence of postoperative high-frequency sensorineural loss in these patients. Using two types of prostheses we found a significant amelioration of speech intelligibility in quiet for the Schuknecht minihole prosthesis as compared to the House wire loop.
A new method for automatic voice-quality registration is presented. The method is based on a technique called phonetography, which is the registration of the dynamic range of a voice as a function of fundamental frequency. In the new phonetogram-recording method fundamental frequency (Fo) and sound-pressure level (SPL) are automatically measured and represented in an XY-diagram. Three additional acoustical voice-quality parameters are measured simultaneously with Fo and SPL: (a) jitter in the Fo as a measure for roughness, (b) the SPL difference between the 0-1.5 kHz and the 1.5-5 kHz bands as a measure for sharpness, and (c) the vocal-noise level above 5 kHz as a measure for breathiness. With this method, the voice-quality parameter values, which may change substantially as a function of Fo and SPL, are pinned to a reference position in the patient's total vocal range. Seen as a reference tool, the phonetogram opens the possibility for a more meaningful comparison of voice-quality data. Some examples, demonstrating the dependence of the chosen quality parameters on Fo and SPL are given.
A study was made of the effect of interaural time delay (ITD) and acoustic headshadow on binaural speech intelligibility in noise. A free-field condition was simulated by presenting recordings, made with a KEMAR manikin in an anechoic room, through earphones. Recordings were made of speech, reproduced in front of the manikin, and of noise, emanating from seven angles in the azimuthal plane, ranging from 0 degree (frontal) to 180 degrees in steps of 30 degrees. From this noise, two signals were derived, one containing only ITD, the other containing only interaural level differences (ILD) due to headshadow. Using this material, speech reception thresholds (SRT) for sentences in noise were determined for a group of normal-hearing subjects. Results show that (1) for noise azimuths between 30 degrees and 150 degrees, the gain due to ITD lies between 3.9 and 5.1 dB, while the gain due to ILD ranges from 3.5 to 7.8 dB, and (2) ILD decreases the effectiveness of binaural unmasking due to ITD (on the average, the threshold shift drops from 4.6 to 2.6 dB). In a second experiment, also conducted with normal-hearing subjects, similar stimuli were used, but now presented monaurally or with an overall 20-dB attenuation in one channel, in order to simulate hearing loss. In addition, SRTs were determined for noise with fixed ITDs, for comparison with the results obtained with head-induced (frequency dependent) ITDs.(ABSTRACT TRUNCATED AT 250 WORDS)