Changes in vowel quality in adult cochlear implant users.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to A J Bosman.
Explore the source record for details and available documents.
A model is presented that quantifies the effect of context on speech recognition. In this model, a speech stimulus is considered as a concatenation of a number of equivalent elements (e.g., phonemes constituting a word). The model employs probabilities that individual elements are recognized and chances that missed elements are guessed using contextual information. Predictions are given of the probability that the entire stimulus, or part of it, is reproduced correctly. The model can be applied to both speech recognition and visual recognition of printed text. It has been verified with data obtained with syllables of the consonant-vowel-consonant (CVC) type presented near the reception threshold in quiet and in noise, with the results of an experiment using orthographic presentation of incomplete CVC syllables and with results of word counts in a CVC lexicon. A remarkable outcome of the analysis is that the cues which occur only in spoken language (e.g., coarticulatory cues) seem to have a much greater influence on recognition performance when the stimuli are presented near the threshold in noise than when they are presented near the absolute threshold. Demonstrations are given of further predictions provided by the model: word recognition as a function of signal-to-noise ratio, closed-set word recognition, recognition of interrupted speech, and sentence recognition.
Successful rehabilitation of the profoundly hearing impaired by means of a speech-processing hearing aid requires integration of auditory and visual speech information. In two studies we investigated (1) which perceptual dimensions play a role in processing various auditory patterns by profoundly hearing-impaired subjects, and (2) which Dutch consonants and vowels can be identified by the average lipreader. One of the important cues in auditory pattern discrimination seems to be the presence of temporal fluctuations (beats) in the signal, resulting from two closely placed frequency components. However, this feature is confounded with the perception of loudness. A second cue used by some subjects is the presence of high-frequency peaks. In lipreading, at least three groups of consonants and three vowel groups may be distinguished; phonemes within a group cannot be discriminated from each other. Important features for both consonants and vowels are degree of lip opening and lip activity (movement or rounding). These results suggest how the auditory speech signal might be coded so as to provide supplementary information to speechreading.
In providing profoundly hearing-impaired persons with processed speech through a signal-processing hearing aid, it is important that the new speech code matches their auditory capacities. This processing capacity for auditory information was investigated in this study. In part 1, the subjects' ability to judge similarities among 8 different but related harmonic complexes was studied. The patterns contained different numbers of harmonics to a 125-Hz fundamental frequency; the harmonics had been spread over the spectrum in various ways. The perceptual judgments appeared to be based on a temporal cue, beat strength, and a spectral cue, related to the balance of high and low frequency components. In part 2, three sets of synthetic vowels were presented to the subjects. Each vowel was realized by summing harmonically related in-phase sinusoids at two formant frequencies. The sets differed in the number of sinusoids per formant: 1, 2 or 3. It was found that the subjects used spectral cues and vowel length for differentiating among the vowels. The overall results show the limited but perhaps usable ability of the profoundly impaired ear to handle spectral information. Implications of these results for the development of signal-processing hearing aids for the profoundly hearing impaired are discussed.
Both meaningful (sense) and meaningless (nonsense) syllables of the consonant-vowel-consonant type (CVC syllables) and short sentences consisting of 8 or 9 syllables were presented in quiet and in noise to 20 young subjects with normal hearing and to three groups of 20 subjects each with presbycusis, with Menière's disease and with noise-induced hearing loss. All materials were uttered by a female speaker. The masking noise consisted of continuous noise shaped in accordance with the long-term average spectrum of the speaker. For each individual, the level of the noise was chosen halfway between the speech reception threshold (SRT) for sentences in quiet and 100 dBA. For all groups of subjects in quiet, the SRT for whole-sentence correct scores (sentence SRT) corresponded closely to the SRT for phoneme scores with sense CVC syllables in quiet (CVC phoneme SRT). Averaged across all groups of subjects, sentence SRT in quiet could be predicted within 4.2 dB from CVC phoneme SRT in quiet and sentence SRT in noise within 1.8 dB from CVC phoneme SRT in noise. The prediction error for sentence SRT in quiet using the pure-tone average (PTA) of 0.5, 1 and 2 kHz was 6.0 dB; for sentence SRT in noise using the PTA of 2 and 4 kHz, it was 2.1 dB. In view of the smaller measurement error, a direct measurement of sentence SRT in noise is advisable.
For many profoundly hearing-impaired listeners (hearing loss > 90 dB HL) speechreading is the most important means of communication; amplified speech may provide, at best, additional information to speechreading. In order to improve audiovisual communication, three speech pattern elements comprising voice-fundamental frequency (f0), the first formant (F1), and the first and the second formant (F1F2) were presented as supplements to speechreading. A fourth condition consisted of a natural speech supplement, a fifth of speechreading only. Twenty subjects were tested; all audiovisual speech scores were significantly higher than the purely visual scores. Audiovisual scores for amplified, natural speech were significantly higher than those for f0 and F1F2 coded speech. Scores for natural speech and for F1 coded speech were not significantly different. The relations between the increase in audiovisual speech scores over the visual scores and measures of difference limen for frequency (DLf) and gap detection were not clear. The most prominent correlations with the speech scores were found for the DLf at 125 Hz and for gap detection.
The present study addresses the effect of cochlear implantation on vowel production of 20 post-lingually deafened Dutch subjects. All subjects received the Nucleus 22 implant (3 WSP and 17 MSP processors). Speech recordings were made pre-implantation and three and twelve months post-implantation with the implant switched on and off. The first and second formant frequencies were measured for eleven Dutch vowels (monophthongs only) in an h-vowel-t context. Twelve months post-implantation, the results showed an increase in the ranges of the first and second formant frequency covered by the respective vowels when the implant was switched on. The increase in the formant frequency range was most marked for some subjects with a relatively small formant range pre-implantation. Also, at 12 months post-implantation with the implant switched on we found a significant shift of the first and second formant frequency towards the normative values. Moreover, at this time the results showed significantly increased clustering of the respective vowels, suggesting an improvement in the ability to produce phonological contrasts between vowels. Clustering is defined as the ratio of the between-vowel variance of the first and second formant frequency and the within-vowel variance of three tokens of the same vowel.
The present study addresses the effect of cochlear implantation on the voice fundamental frequency at which 20 post-lingually deafened Dutch subjects utter speech materials. All subjects received the Nucleus 22 cochlear implant (3 WSP and 17 MSP processors). Speech recordings were made pre-implantation and three and twelve months post-implantation with the implant switched on and off. The fundamental frequency (f0) was sampled while reading a text. The pre-implantation results show that in some subjects, f0 was too high compared with the range in f0 of normally-hearing subjects. Post-implantation, with the implant switched on, we found that the abnormally high f0 values pre-implantation changed toward the normative values. In addition, post-implantation we found that the range over which f0 varied within a subject while reading the text, the f0 sway, decreased for most subjects who, pre-implantation, had their f0 sway outside the normative ranges, the normative range being defined as the interval between the mean plus/minus one standard deviation of the f0 sway found for normally-hearing subjects. Voice fundamental frequency of post-lingually deafened adults is characterized by large interindividual variability in the pre-implantation f0 values. This large interindividual variability is found also in the effect of cochlear implantation on f0.
Auditory alone, visual alone and audiovisual recognition of consonant-vowel consonant syllables were measured in 32 severely hearing-impaired children with hearing loss (PTA) in a narrow range around 90 dB HL when using their hearing aids. Multidimensional scaling analysis (INDSCAL) and information transmission analysis (ITA), applied to the confusion matrices obtained from the responses in each presentation mode and for each phoneme category, revealed perceptual dimensions and percentages of transmitted feature information (PTI). These were studied in relation to PTA, the auditory alone score and in relation to the efficiency of the audiovisual interaction (enhancement) over the probalistic summation of the auditory alone and visual alone score. INDSCAL analysis shows that auditory alone recognition of vowels is based on the perceptual dimensions F2 and F1 and that of consonants on the dimensions 'frication' and 'voicing'. In the auditory mode the interpretation of the INDSCAL dimensions in the stimulus spaces is in reasonable agreement with the ITA results. PTI decreases gradually with decreasing auditory alone phoneme score. Audiovisual recognition of vowels is based on a combination of the auditory dimension 'open/closed' (F1), and the visual dimensions 'lip rounding' and 'vertical lip opening'. Audiovisual recognition of initial consonants is based on a combination of the visual dimension 'front/back' and the auditory dimension 'continuance'. Recognition of final consonants is based on a combination of the visual dimension 'front/back' and an uninterpretable dimension. The perceptual dimensions are independent of both the level of the auditory alone phoneme score and audiovisual enhancement. Audiovisual enhancement is mainly a property of an individual and independent of both auditory alone and visual alone scores. ITA analysis, based on a phonological classification of the features, supports the results of the INDSCAL analysis in the auditory alone mode. It is not useful in the description of the audiovisual interaction, probably due to the phonological basis of the feature classification.
The present study addresses the effect of cochlear implantation on the intelligibility of vowels produced by 20 post-lingually deafened Dutch subjects. All subjects received the Nucleus-22 cochlear implant (3 WSP and 17 MSP processors). Speech recordings were made pre-implantation and three and twelve months post-implantation with the implant switched on and off. Vowel intelligibility (monophthongs only) was determined using a panel of listeners. For all implanted subjects intelligibility was measured in a noisy background. For seven poorly speaking subjects it was also measured in a quiet background. After implantation with the Nucleus-22 device the results showed that vowel intelligibility, measured for all subjects in a noisy background, increased for most of them (about 15), while it increased for about half the number of poorly speaking subjects measured in a quiet background. Twelve months after implantation vowel intelligibility, measured for all subjects in noise, appeared to be based on first and second formant information. This was also found for the subgroup of seven subjects performing poorly pre-implantation when analysed separately. However, vowel intelligibility for this subgroup, when measured in a quiet background, was based also on vowel duration. The differences between the overall result in noise and the results of the subgroup in quiet should be attributed mainly to the noise and not to aspects of poor speech production in the subgroup. In addition, this study addresses the relationship between the intelligibility scores and objective measurements of vowel quality performed in a previous study. The results showed that the vowel intelligibility scores are mainly determined by the position of the second formant frequencies.