PubMed HealthSearch

Biomedical subjects

L D Braida

Publications and source records attributed to L D Braida.

At least 19 recordsLinked to original sources

Analytic study of the Tadoma method: improving performance through the use of supplementary tactual displays.

Although results obtained with the Tadoma method of speechreading have set a new standard for tactual speech communication, they are nevertheless inferior to those obtained in the normal auditory domain. Speech reception through Tadoma is comparable to that of normal-hearing subjects listening to speech under adverse conditions corresponding to a speech-to-noise ratio of roughly 0 dB. The goal of the current study was to demonstrate improvements to speech reception through Tadoma through the use of supplementary tactual information, thus leading to a new standard of performance in the tactual domain. Three supplementary tactual displays were investigated: (a) an articulatory-based display of tongue contact with the hard palate; (b) a multichannel display of the short-term speech spectrum; and (c) tactual reception of Cued Speech. The ability of laboratory-trained subjects to discriminate pairs of speech segments that are highly confused through Tadoma was studied for each of these augmental displays. Generally, discrimination tests were conducted for Tadoma alone, the supplementary display alone, and Tadoma combined with the supplementary tactual display. The results indicated that the tongue-palate contact display was an effective supplement to Tadoma for improving discrimination of consonants, but that neither the tongue-palate contact display nor the short-term spectral display was highly effective in improving vowel discriminability. For both vowel and consonant stimulus pairs, discriminability was nearly perfect for the tactual reception of the manual cues associated with Cued Speech. Further experiments on the identification of speech segments were conducted for Tadoma combined with Cued Speech. The observed data for both discrimination and identification experiments are compared with the predictions of models of integration of information from separate sources.

Blindness

Single band amplitude envelope cues as an aid to speechreading.

Amplitude envelopes derived from speech have been shown to facilitate speech-reading to varying degrees, depending on how the envelope signals were extracted and presented and on the amount of training given to the subjects. In this study, three parameters related to envelope extraction and presentation were examined using both easy and difficult sentence materials: (1) the bandwidth and centre frequency of the filtered speech signal used to obtain the envelope; (2) the bandwidth of the envelope signal determined by the lowpass filter cutoff frequency used to "smooth" the envelope fluctuations; and (3) the carrier signal used to convey the envelope cues. Results for normal hearing subjects following a brief visual and auditory-visual familiarization/training period showed that (1) the envelope derived from wideband speech does not provide the greatest benefit to speechreading when compared to envelopes derived from selected octave bands of speech; (2) as the bandwidth centred around the carrier frequency increased from 12.5 to 1600 Hz, auditory-visual (AV) performance obtained with difficult sentence materials improved, especially for envelopes derived from high-frequency speech energy; (3) envelope bandwidths below 25 Hz resulted in AV scores that were sometimes equal to or worse than speechreading alone; (4) for each filtering condition tested, there was at least one bandwidth and carrier condition that produced AV scores that were significantly greater than speechreading alone; (5) low-frequency carriers were better than high-frequency or wideband carriers for envelopes derived from an octave band of speech centred at 500 Hz; and (6) low-frequency carriers were worse than high-frequency or wideband carriers for envelopes derived from an octave band centred at 3150 Hz. These results suggest that amplitude envelope cues can provide a substantial benefit to speechreading for both easy and difficult sentence materials, but that frequency transposition of these signals to regions remote from their "natural" spectral locations may result in reduced performance.

Adult

Crossmodal integration in the identification of consonant segments.

Although speechreading can be facilitated by auditory or tactile supplements, the process that integrates cues across modalities is not well understood. This paper describes two "optimal processing" models for the types of integration that can be used in speechreading consonant segments and compares their predictions with those of the Fuzzy Logical Model of Perception (FLMP, Massaro, 1987). In "pre-labelling" integration, continuous sensory data is combined across modalities before response labels are assigned. In "post-labelling" integration, the responses that would be made under unimodal conditions are combined, and a joint response is derived from the pair. To describe pre-labelling integration, confusion matrices are characterized by a multidimensional decision model that allows performance to be described by a subject's sensitivity and bias in using continuous-valued cues. The cue space is characterized by the locations of stimulus and response centres. The distance between a pair of stimulus centres determines how well two stimuli can be distinguished in a given experiment. In the multimodal case, the cue space is assumed to be the product space of the cue spaces corresponding to the stimulation modes. Measurements of multimodal accuracy in five modern studies of consonant identification are more consistent with the predictions of the pre-labelling integration model than the FLMP or the post-labelling model.

Attention

Evaluating the articulation index for auditory-visual input.

An investigation of the auditory-visual (AV) articulation index (AI) correction procedure outlined in the ANSI standard [ANSI S3.5-1969 (R1986)] was made by evaluating auditory (A), visual (V), and auditory-visual sentence identification for both wideband speech degraded by additive noise and a variety of bandpass-filtered speech conditions presented in quiet and in noise. When the data for each of the different listening conditions were averaged across talkers and subjects, the procedure outlined in the standard was fairly well supported, although deviations from the predicted AV score were noted for individual subjects as well as individual talkers. For filtered speech signals with AIA less than 0.25, there was a tendency for the standard to underpredict AV scores. Conversely, for signals with AIA greater than 0.25, the standard consistently overpredicted AV scores. Additionally, synergistic effects, where the AIA obtained from the combination of different bandpass-filtered conditions was greater than the sum of the individual AIA's, were observed for all nonadjacent filter-band combinations (e.g., the addition of a low-pass band with a 630-Hz cutoff and a high-pass band with a 3150-Hz cutoff). These latter deviations from the standard violate the basic assumption of additivity stated by Articulation Theory, but are consistent with earlier reports by Pollack [I. Pollack, J. Acoust. Soc. Am. 20, 259-266 (1948)], Licklider [J. C. R. Licklider, Psychology: A Study of a Science, Vol. 1, edited by S. Koch (McGraw-Hill, New York, 1959), pp. 41-144], and Kryter [K. D. Kryter, J. Acoust. Soc. Am. 32, 547-556 (1960)].

Adult

Development and testing of artificial low-frequency speech codes.

In a new approach to the frequency-lowering of speech, artificial codes were developed for 24 consonants (C) and 15 vowels (V) for two values of lowpass cutoff frequency F (300 and 500 Hz). Each individual phoneme was coded by a unique, nonvarying acoustic signal confined to frequencies less than or equal to F. Stimuli were created through variations in spectral content, amplitude, and duration of tonal complexes or bandpass noise. For example, plosive and fricative sounds were constructed by specifying the duration and relative amplitude of bandpass noise with various center frequencies and bandwidths, while vowels were generated through variations in the spectral shape and duration of a ten-tone harmonic complex. The ability of normal-hearing listeners to identify coded Cs and Vs in fixed-context syllables was compared to their performance on single-token sets of natural speech utterances lowpass filtered to equivalent values of F. For a set of 24 consonants in C-/a/ context, asymptotic performance on coded sounds averaged 90 percent correct for F = 500 Hz and 65 percent for F = 300 Hz, compared to 75 percent and 40 percent for lowpass filtered speech. For a set of 15 vowels in /b/-V-/t/ context, asymptotic performance on coded sounds averaged 85 percent correct for F = 500 Hz and 65 percent for F = 300 Hz, compared to 85 percent and 50 percent for lowpass filtered speech. Identification of coded signals for F = 500 Hz was also examined in CV syllables where C was selected at random from the set of 24 Cs and V was selected at random from the set of 15 Vs. Asymptotic performance of roughly 67 percent correct and 71 percent correct was obtained for C and V identification, respectively. These scores are somewhat lower than those obtained in the fixed-context experiments. Finally, results were obtained concerning the effect of token variability on the identification of lowpass filtered speech. These results indicate a systematic decrease in percent-correct score as the number of tokens representing each phoneme in the identification tests increased from one to nine.

Evaluation Studies as Topic

Analytic study of the Tadoma method: effects of hand position on segmental speech perception.

In the Tadoma method of communication, deaf-blind individuals receive speech by placing a hand on the face and neck of the talker and monitoring actions associated with speech production. Previous research has documented the speech perception, speech production, and linguistic abilities of highly experienced users of the Tadoma method. The current study was performed to gain further insight into the cues involved in the perception of speech segments through Tadoma. Small-set segmental identification experiments were conducted in which the subjects' access to various types of articulatory information was systematically varied by imposing limitations on the contact of the hand with the face. Results obtained on 3 deaf-blind, highly experienced users of Tadoma were examined in terms of percent-correct scores, information transfer, and reception of speech features for each of sixteen experimental conditions. The results were generally consistent with expectations based on the speech cues assumed to be available in the various hand positions.

Adult

Speaking clearly for the hard of hearing. III: An attempt to determine the contribution of speaking rate to differences in intelligibility between clear and conversational speech.

Previous studies (Picheny, Durlach, & Braida, 1985, 1986) have demonstrated that substantial intelligibility differences exist for hearing-impaired listeners for speech spoken clearly compared to speech spoken conversationally. This paper presents the results of a probe experiment intended to determine the contribution of speaking rate to the intelligibility differences. Clear sentences were processed to have the durational properties of conversational speech, and conversational sentences were processed to have the durational properties of clear speech. Intelligibility testing with hearing-impaired listeners revealed both sets of materials to be degraded after processing. However, the degradation could not be attributable to processing artifacts because reprocessing the materials to restore their original durations produced intelligibility scores close to those observed for the unprocessed materials. We conclude that the simple processing to alter the relative durations of the speech materials was not adequate to assess the contribution of speaking rate to the intelligibility differences; further studies are proposed to address this question.

Hearing Loss, Sensorineural

Tactile communication of speech: comparison of two computer-based displays.

Two methods of encoding speech for tactile displays were compared in discrimination experiments using speech segments. One display represented the short-term speech spectrum in time-swept mode and used vibration amplitude to encode spectral amplitude. The other represented the linear predictive coding (LPC)-derived vocal tract shape as a filled bar graph in which the number of active vibrators was used to encode cross sectional area. The displays were applied to the thigh via a matrix of vibrators. The vibrators were driven at 250 Hz during voiced segments, and by random noise during unvoiced segments. Overall results show a slight superiority for the spectral display in vowel discrimination. Detailed results were analyzed in terms of an articulatory description of the speech stimuli, a multidimensional scaling (MDS) analysis of confusions, and an ideal receiver analysis. The results of these analyses suggest that the detailed characteristics of the tactile patterns were only crudely discriminated.

Computer Systems

Spectral-shape discrimination. I. Results from normal-hearing listeners for stationary broadband noises.

This research is concerned with the ability of normal-hearing listeners to discriminate broadband signals on the basis of spectral shape. The signals were six broadband noises whose spectral shapes were modeled after the spectra of unvoiced fricative and plosive consonants. The difficulty of the discriminations was controlled by the addition of noise filtered to match the long-term speech spectrum. Two-interval discrimination measurements were made in which loudness cues were eliminated by randomizing (roving) the overall stimulus level between presentation intervals. Experimental results, examined as a function of intensity rove width, stimulus duration, and stimulus pair, were related to the predictions of a simple filter-bank model whose fitting parameter provides an estimate of internal noise. Most results, with the notable exception of duration effects, were predicted by the model. Estimates of internal noise in each frequency channel averaged roughly 7 dB for long-duration stimuli and 13 dB for short-duration stimuli. Results and predictions are compared to results of other studies concerned with the discrimination of spectral shape.

Humans

Principal-component amplitude compression for the hearing impaired.

Principal-component amplitude compression, a means for matching speech to the reduced dynamic range in sensorineural hearing impairments, is a multiband approach aimed at preserving details of spectral shape while reducing overall level variation. The effect of compression has been studied for the first and second principal components (PC1 an PC2) of the short-term speech spectrum, which are roughly representative of overall level and spectral tilt, respectively. Compression of PC1 roughly equalizes consonant and vowel levels while compression of PC2 provides time-varying high-frequency emphasis. The effect on speech intelligibility of sensorineural hearing-impaired listeners of two principal-component compression system implementations, compression of PC1 and compression of both PC1 and PC2, was compared to that of linear amplification (LA), independent compression of multiple bands (MBC), and wideband compression (WC). Results indicate that compression of overall level as provided by compression of PC1 and WC improved intelligibility relative to LA over a 10- to 15-dB range of input levels. While MBC was beneficial in some cases, it did not provide higher intelligibility than WC. Compression of PC2 did not benefit but rather degraded performance relative to LA. Error analyses and band-level measurements indicate that the highest intelligibility is obtained when audibility is improved and the relative spectral shapes of different speech sounds are preserved.

Auditory Threshold

Multiband compression limiting for hearing-impaired listeners.

Four multiband compression limiters and two linear amplification systems were compared in terms of the intelligibility of consonant-vowel-consonant (CVC) nonsense syllables for two hearing-impaired listeners over a 30 dB range of input levels. Each system incorporated one of two frequency-gain characteristics and one of three limiting characteristics (no limiting, moderate limiting, or severe limiting). The subjects were instructed to choose overall listening levels that would permit speech spanning the range of input levels to be as intelligible as possible and comfortable for long-term listening. Relative to linear amplification, the overall gain selected by the subjects increased by roughly 5 and 11 dB for the moderate and severe limiter, respectively. With linear amplification, the maximum score, 82 percent correct, was obtained at the highest input level and scores fell roughly 34 percentage points as input level was reduced. With compression limiting, although the maximum scores, 81 percent and 79 percent correct, were obtained at lower input levels, performance was comparable to that with linear amplification. Also, scores spanned a range of only 22 and 9 percentage points across the range of input levels with the moderate and severe limiter, respectively. This benefit was due to the improved scores provided by compression limiting at the low input levels. However, this advantage was offset somewhat by the disadvantage provided by compression at high input levels relative to linear amplification. Error analysis indicated that the spectral degradations introduced by independent compression of 16 frequency bands may have caused the reduced intelligibility at higher input levels.

Adult

Speaking clearly for the hard of hearing. II: Acoustic characteristics of clear and conversational speech.

The first paper of this series (Picheny, Durlach, & Braida, 1985) presented evidence that there are substantial intelligibility differences for hearing-impaired listeners between nonsense sentences spoken in a conversational manner and spoken with the effort to produce clear speech. In this paper, we report the results of acoustic analyses performed on the conversational and clear speech. Among these results are the following. First, speaking rate decreases substantially in clear speech. This decrease is achieved both by inserting pauses between words and by lengthening the durations of individual speech sounds. Second, there are differences between the two speaking modes in the numbers and types of phonological phenomena observed. In conversational speech, vowels are modified or reduced, and word-final stop bursts are often not released. In clear speech, vowels are modified to a lesser extent, and stop bursts, as well as essentially all word-final consonants, are released. Third, the RMS intensities for obstruent sounds, particularly stop consonants, is greater in clear speech than in conversational speech. Finally, changes in the long-term spectrum are small. Thus, speaking clearly cannot be regarded as equivalent to the application of high-frequency emphasis.

Hearing Loss, Sensorineural

Towards a model for discrimination of broadband signals.

The conventional model for broadband discrimination assumes that resolution is limited by peripheral internal noise that is statistically independent across channels. In this paper, we extend this model in a number of directions. In particular, we compute, compare, and discuss the effects of interchannel correlation and central noise on the sensitivity index d', for discrimination of overall level and discrimination of spectral shape.

Auditory Perception

Multichannel syllabic compression for severely impaired listeners.

Two listeners with congenital hearing losses characterized by flat audiograms and dynamic ranges of 18-33 dB were tested with three compression systems and one (reference) linear amplification system. The compression systems placed progressively larger amounts of speech energy within the listener's residual dynamic range, by raising to audibility and compressing 25, 50, and 90 percent of the short-term input amplitude distribution in each of 16 frequency bands. The comparison linear system was defined by adjusting six octave-wide bands of speech to comfortable levels. System performance was evaluated with nonsence CVC syllables presented at a constant input level and spoken by two talkers. Extensive training was provided to ensure stable performance. The results were notably speaker-dependent, with compression consistently providing better performance for one speaker, linear amplification for the other. Averaged over speakers, however, there was no net advantage for any of the compression systems for any listener. The use of high compression ratios and large input ranges tended to degrade perception of initial consonants and vowels. Under some conditions, however, final consonant scores were higher with compression than with linear amplification. Compression generally enhanced the distinction between stops and fricatives, but degraded spectral-concentration and relative-intensity cues required to identify place of articulation.

Adult

Speaking clearly for the hard of hearing I: Intelligibility differences between clear and conversational speech.

This paper is concerned with variations in the intelligibility of speech produced for hearing-impaired listeners under two conditions. Estimates were made of the magnitude of the intelligibility differences between attempts to speak clearly and attempts to speak conversationally. Five listeners with sensorineural hearing losses were tested on groups of nonsense sentences spoken clearly and conversationally by three male talkers as a function of level and frequency-gain characteristic. The average intelligibility difference between clear and conversational speech averaged across talker was found to be 17 percentage points. To a first approximation, this difference was independent of the listener, level, and frequency-gain characteristic. Analysis of segmental-level errors was only possible for two listeners and indicated that improvements in intelligibility occurred across all phoneme classes.

Adult

Research on the Tadoma method of speech communication.

In Tadoma, speech is received by placing a hand on the talker's face and monitoring actions associated with speech production. Our initial research has documented the speech perception, speech production, and linguistic abilities of deaf-blind individuals highly trained in Tadoma. This research has demonstrated that good speech reception can be achieved through the tactile sense: Performance is roughly equivalent to that of normals listening in noise or babble with a signal-to-noise ratio in the range 0-6 dB. It appears that the principal cues employed are lip movement, jaw movement, oral airflow, and laryngeal vibration, and that the errors which occur are caused primarily by inadequate information on tongue position. Our current research includes (1) learning of Tadoma by normal subjects with simulated deafness and blindness, (2) augmenting Tadoma with a supplemental tactile display of tongue position, and (3) developing a synthetic Tadoma system in which signals recorded from a talker's face are used to drive an artificial face. This research is expected to increase our understanding of Tadoma and its relation to other tactile communication methods, show that performance obtained through Tadoma does not represent the ultimate limits of the tactile sense, and provide a research tool for studying transformations of Tadoma.

Adult

Discrimination and identification of frequency-lowered speech in listeners with high-frequency hearing impairment.

The effects of frequency lowering on consonant perception were studied in four listeners with high-frequency sensorineural loss. Frequency lowering, accomplished by pitch-invariant nonuniform compression of the short-term spectral envelope, included lowering to bandwidths of 2500 and 1250 Hz. Performance on frequency lowering was compared to that obtained with linear amplification using uniform or high-frequency emphasis. Results of pairwise discrimination tests indicated that performance on lowering to 1250 Hz was inferior to that obtained with linear amplification. After training, performance on consonant identification with lowering to bandwidths of 2500 or 1250 Hz was equivalent or inferior to that obtained with linear amplification, depending on the subject. In most cases, the performance of the impaired subjects on a given lowering condition was inferior to that obtained by normal subjects, except for one impaired subject whose performance on two lowering conditions was similar to normal.

Hearing Loss, Sensorineural