PubMed HealthSearch

Biomedical subjects

N I Durlach

Publications and source records attributed to N I Durlach.

At least 19 recordsLinked to original sources

Analytic study of the Tadoma method: improving performance through the use of supplementary tactual displays.

Although results obtained with the Tadoma method of speechreading have set a new standard for tactual speech communication, they are nevertheless inferior to those obtained in the normal auditory domain. Speech reception through Tadoma is comparable to that of normal-hearing subjects listening to speech under adverse conditions corresponding to a speech-to-noise ratio of roughly 0 dB. The goal of the current study was to demonstrate improvements to speech reception through Tadoma through the use of supplementary tactual information, thus leading to a new standard of performance in the tactual domain. Three supplementary tactual displays were investigated: (a) an articulatory-based display of tongue contact with the hard palate; (b) a multichannel display of the short-term speech spectrum; and (c) tactual reception of Cued Speech. The ability of laboratory-trained subjects to discriminate pairs of speech segments that are highly confused through Tadoma was studied for each of these augmental displays. Generally, discrimination tests were conducted for Tadoma alone, the supplementary display alone, and Tadoma combined with the supplementary tactual display. The results indicated that the tongue-palate contact display was an effective supplement to Tadoma for improving discrimination of consonants, but that neither the tongue-palate contact display nor the short-term spectral display was highly effective in improving vowel discriminability. For both vowel and consonant stimulus pairs, discriminability was nearly perfect for the tactual reception of the manual cues associated with Cued Speech. Further experiments on the identification of speech segments were conducted for Tadoma combined with Cued Speech. The observed data for both discrimination and identification experiments are compared with the predictions of models of integration of information from separate sources.

Blindness

Development and testing of artificial low-frequency speech codes.

In a new approach to the frequency-lowering of speech, artificial codes were developed for 24 consonants (C) and 15 vowels (V) for two values of lowpass cutoff frequency F (300 and 500 Hz). Each individual phoneme was coded by a unique, nonvarying acoustic signal confined to frequencies less than or equal to F. Stimuli were created through variations in spectral content, amplitude, and duration of tonal complexes or bandpass noise. For example, plosive and fricative sounds were constructed by specifying the duration and relative amplitude of bandpass noise with various center frequencies and bandwidths, while vowels were generated through variations in the spectral shape and duration of a ten-tone harmonic complex. The ability of normal-hearing listeners to identify coded Cs and Vs in fixed-context syllables was compared to their performance on single-token sets of natural speech utterances lowpass filtered to equivalent values of F. For a set of 24 consonants in C-/a/ context, asymptotic performance on coded sounds averaged 90 percent correct for F = 500 Hz and 65 percent for F = 300 Hz, compared to 75 percent and 40 percent for lowpass filtered speech. For a set of 15 vowels in /b/-V-/t/ context, asymptotic performance on coded sounds averaged 85 percent correct for F = 500 Hz and 65 percent for F = 300 Hz, compared to 85 percent and 50 percent for lowpass filtered speech. Identification of coded signals for F = 500 Hz was also examined in CV syllables where C was selected at random from the set of 24 Cs and V was selected at random from the set of 15 Vs. Asymptotic performance of roughly 67 percent correct and 71 percent correct was obtained for C and V identification, respectively. These scores are somewhat lower than those obtained in the fixed-context experiments. Finally, results were obtained concerning the effect of token variability on the identification of lowpass filtered speech. These results indicate a systematic decrease in percent-correct score as the number of tokens representing each phoneme in the identification tests increased from one to nine.

Evaluation Studies as Topic

Manual discrimination of force using active finger motion.

In these experiments, two plates were grasped between the thumb and forefinger and squeezed together along a linear track. An electromechanical system presented a constant resistance force during the squeeze up to a predetermined location on the track, whereupon the force effectively went to infinity (simulating a wall) or to zero (simulating a cliff). The task of the subject was to discriminate between two alternative levels of the constant resistance force (a reference level and a reference-plus-increment level). Results of these experiments indicate a just noticeable difference of roughly 7% of the reference force using a one-interval paradigm with trial-by-trial feedback over the ranges 2.5 less than or equal to F0 less than or equal to 10.0 newtons, 5 less than or equal to D less than or equal to 30 mm, 45 less than or equal to S less than or equal to 125 mm, and 25 less than or equal to V less than or equal to 160 mm/sec, where F0 is the reference force, D is the distance squeezed, S is the initial fingerspan, and V is the mean velocity of the squeeze. These results, based on tests with 5 subjects, are consistent with a wide range of previous results, some of which are associated with other body surfaces and muscle systems and many of which were obtained with different psychophysical methods.

Biomechanical Phenomena

A study of the tactual and visual reception of fingerspelling.

A method of communication in frequent use among members of the deaf-blind community is the tactual reception of fingerspelling. In this method, the hand of the deaf-blind individual is placed on the hand of the sender to monitor the handshapes and movements associated with the letters of the manual alphabet. The purpose of the current study was to examine the ability of experienced deaf-blind subjects to receive fingerspelled materials, including sentences and connected text, through the tactual sense. A parallel study of the reception of fingerspelling through the visual sense was also conducted using sighted deaf subjects. For both visual and tactual reception of fingerspelled sentences, accuracy of reception was examined as a function of rate of presentation. In the tactual study, where rates were limited to those that could be produced naturally by an experienced interpreter, highly accurate reception of conversational sentence materials was observed throughout the range of naturally produced rates (i.e., 2 to 6 letters/s). In the visual study, rates in excess of those that can be produced naturally were achieved through variable-speed playback of videotapes of fingerspelled sentences. The results of this study indicate that performance varies systematically as a function of rate of presentation, with scores of 50% correct on conversational sentences obtained at rates of 12 to 16 letters/s (i.e., rates roughly double to triple normal speed). These results suggest that normal communication rates for the visual reception of fingerspelling are restricted by limitations on the rate of manual production. Although maximal rates of natural manual production of fingerspelling correspond to the presentation of a new handshape on the order of once every 150-20 ms, the data from the sped-up visual study suggest that experienced receivers of visual fingerspelling are able to receive sentences at substantially higher rates of fingerspelling (which are, in fact, comparable to communication rates for spoken English).

Adult

Analytic study of the Tadoma method: effects of hand position on segmental speech perception.

In the Tadoma method of communication, deaf-blind individuals receive speech by placing a hand on the face and neck of the talker and monitoring actions associated with speech production. Previous research has documented the speech perception, speech production, and linguistic abilities of highly experienced users of the Tadoma method. The current study was performed to gain further insight into the cues involved in the perception of speech segments through Tadoma. Small-set segmental identification experiments were conducted in which the subjects' access to various types of articulatory information was systematically varied by imposing limitations on the contact of the hand with the face. Results obtained on 3 deaf-blind, highly experienced users of Tadoma were examined in terms of percent-correct scores, information transfer, and reception of speech features for each of sixteen experimental conditions. The results were generally consistent with expectations based on the speech cues assumed to be available in the various hand positions.

Adult

Range effects in the identification of lateral position.

Experiments on the identification of interaural time and interaural amplitude differences were conducted to evaluate the effects of stimulus range on identification performance. Three stimulus sets, large range (LR), small-range center (SRC), and small-range side (SRS), were used in experiments on interaural time and amplitude identification. As expected, data for both sets of measurements show worse resolution for LR than SRC or SRS, demonstrating that the ability to distinguish between two fixed interaural differences can be strongly influenced by the total range of such differences in the stimulus set.

Auditory Perception

Analysis of a synthetic Tadoma system as a multidimensional tactile display.

The Tadoma method is a means of speech reception based on tactile monitoring of the articulatory process. A "synthetic" Tadoma system, involving an artificial face with six facial actions, has been developed as a first-order approximation to the natural Tadoma system. Experiments were conducted to explore the information-transmission characteristics of the synthetic Tadoma system in terms of the four facial movements it incorporates: upper lip in-out, lower lip in-out, lower lip up-down, and jaw up-down movements. Discrimination experiments showed that the just-noticeable difference associated with each movement is about 9% of the reference displacement. One-dimensional (1-D) absolute identification experiments produced, on the average, 1.6 bits of information transfer. Four dimensional (4-D) identification experiments produced information transfers in the range of 3-4 bits. Of the four dimensions considered, performance on the lower lip up-down movement was most affected, and performance on the jaw up-down movement was least affected, by simultaneous roving movements on the other dimensions. As a result of the interaction among the movement channels, the sum of the 1-D information transfers exceeds the 4-D information transfer. However, the sum of the 1-D information transfers obtained from tests with roving parameters is approximately equal to the 4-D information transfer (possibly exemplifying a "generalized information-transfer additivity law"). In general, both the discrimination and identification results appear unexceptional and, hence, the reception of facial movement information by itself does not appear to account for the extraordinary success of the Tadoma method.

Humans

Manual discrimination and identification of length by the finger-span method.

Experiments were conducted on length resolution for objects held between the thumb and fore-finger. The just noticeable difference in length measured in discrimination experiments is roughly 1 mm for reference lengths of 10 to 20 mm. It increases monotonically with reference length but violates Weber's law. Also, it decreases when the subject is permitted to maintain a constant finger span between trials; however, it tends to increase when the nondominant hand is used. As would be expected from studies of other stimulus dimensions in other sense modalities, resolution is considerably poorer in identification experiments than in discrimination experiments. For stimulus sets that cover a broad range (90 mm), the total information transfer is roughly 2 bits; for those that cover a relatively small range (18 mm), it is roughly 1 bit. The data are analyzed and interpreted using analysis techniques and models that have been used previously in studies of audition (e.g., Durlach & Braida, 1969).

Adult

Speaking clearly for the hard of hearing. III: An attempt to determine the contribution of speaking rate to differences in intelligibility between clear and conversational speech.

Previous studies (Picheny, Durlach, & Braida, 1985, 1986) have demonstrated that substantial intelligibility differences exist for hearing-impaired listeners for speech spoken clearly compared to speech spoken conversationally. This paper presents the results of a probe experiment intended to determine the contribution of speaking rate to the intelligibility differences. Clear sentences were processed to have the durational properties of conversational speech, and conversational sentences were processed to have the durational properties of clear speech. Intelligibility testing with hearing-impaired listeners revealed both sets of materials to be degraded after processing. However, the degradation could not be attributable to processing artifacts because reprocessing the materials to restore their original durations produced intelligibility scores close to those observed for the unprocessed materials. We conclude that the simple processing to alter the relative durations of the speech materials was not adequate to assess the contribution of speaking rate to the intelligibility differences; further studies are proposed to address this question.

Hearing Loss, Sensorineural

Tactile communication of speech: comparison of two computer-based displays.

Two methods of encoding speech for tactile displays were compared in discrimination experiments using speech segments. One display represented the short-term speech spectrum in time-swept mode and used vibration amplitude to encode spectral amplitude. The other represented the linear predictive coding (LPC)-derived vocal tract shape as a filled bar graph in which the number of active vibrators was used to encode cross sectional area. The displays were applied to the thigh via a matrix of vibrators. The vibrators were driven at 250 Hz during voiced segments, and by random noise during unvoiced segments. Overall results show a slight superiority for the spectral display in vowel discrimination. Detailed results were analyzed in terms of an articulatory description of the speech stimuli, a multidimensional scaling (MDS) analysis of confusions, and an ideal receiver analysis. The results of these analyses suggest that the detailed characteristics of the tactile patterns were only crudely discriminated.

Computer Systems

Preliminary results of speech-reception tests obtained with the synthetic Tadoma system.

In the Tadoma method of speech reception used by some deaf-blind individuals, speech is understood by placing a hand on the face of the talker and feeling certain mechanical actions of the face associated with speech production. The synthetic Tadoma system is a computer-driven artificial face that simulates these mechanical actions. This paper reports some preliminary data on the discrimination of nonsense syllables with the synthetic system. Although further work is required to produce an accurate simulation, the results suggest that such a goal is indeed achievable.

Blindness

Spectral-shape discrimination. I. Results from normal-hearing listeners for stationary broadband noises.

This research is concerned with the ability of normal-hearing listeners to discriminate broadband signals on the basis of spectral shape. The signals were six broadband noises whose spectral shapes were modeled after the spectra of unvoiced fricative and plosive consonants. The difficulty of the discriminations was controlled by the addition of noise filtered to match the long-term speech spectrum. Two-interval discrimination measurements were made in which loudness cues were eliminated by randomizing (roving) the overall stimulus level between presentation intervals. Experimental results, examined as a function of intensity rove width, stimulus duration, and stimulus pair, were related to the predictions of a simple filter-bank model whose fitting parameter provides an estimate of internal noise. Most results, with the notable exception of duration effects, were predicted by the model. Estimates of internal noise in each frequency channel averaged roughly 7 dB for long-duration stimuli and 13 dB for short-duration stimuli. Results and predictions are compared to results of other studies concerned with the discrimination of spectral shape.

Humans

Masker-bandwidth dependence in homophasic and antiphasic tone detection.

Thresholds for the detection of 500-ms tones in noise at either 0.25 or 4 kHz and either interaurally in-phase (NoSo) or out-of-phase (NoS pi) were measured as a function of the bandwidth of a diotic masking noise (with total noise power held constant). NoSo thresholds followed the classic trend indicative of an effective critical band in noise masking. NoS pi thresholds also indicated critical-band filtering, but with a wider effective critical bandwidth. At subcritical masker bandwidths, for both 0.25 and 4 kHz, NoS pi thresholds increased with an increase in noise bandwidth, despite the fact that total noise power was constant. This latter finding is attributed to binaural insensitivity to rapid fluctuations in the interaural cues that subserve detection in the NoS pi condition. These results suggest a new interpretation for the small difference between NoSo and NoS pi thresholds measured with high-frequency tones in a broadband noise masker. This interpretation is based partly on the inability to utilize rapidly varying interaural cues, partly on out-of-band interference effects, and partly on loss of information related to stimulus fine structure.

Adult

Comparative learning of pitch and loudness identification.

This study investigated possible similarities between the ability to identify pitches and the ability to identify loudnesses. Systematic training of musically naive subjects indicated that frequency identification performance improves at about the same rate as intensity identification performance. Examination of frequency and intensity identification behavior of musically trained subjects showed that their ability to code pitch information efficiently does not generalize to an ability to encode loudness information more efficiently than untrained subjects. Intensity identification training curves of musically trained and untrained subjects are similar, but final performance levels are below frequency identification performance levels exhibited by musically trained subjects, especially those with absolute pitch.

Discrimination Learning

Multidimensional tactile displays: identification of vibratory intensity, frequency, and contactor area.

Experiments were conducted to determine the ability of subjects to identify vibrotactile stimuli presented to the distal pad of the middle finger. The stimulus sets varied along one or more of the following dimensions: intensity of vibration, frequency of vibration, and contactor area. Identification performance was measured by information transfer. One-dimensional stimulus sets produced values in the range 1-2 bits and, for most subjects, three-dimensional sets produced values in the range 4-5 bits. Of the three dimensions considered, performance on the intensity variable was most affected, and performance on contactor area least affected, by simultaneous variations in the other dimensions.

Attention

Multimicrophone adaptive beamforming for interference reduction in hearing aids.

To reduce interference in monaural hearing aids from sound sources that are spatially separated from a target source, we are investigating methods for combining information from multiple microphones. In this paper, we describe an adaptive beamforming method that functions to preserve target signals arriving from straight-ahead of a microphone array while minimizing output power from off-axis interference sources. In a preliminary evaluation of a two-microphone system, sentence intelligibility tests were administered to normal-hearing subjects using processed and unprocessed materials from simulated environments in which the target was on-axis, the interference (speech babble) was 45 degrees off-axis, and the reverberation mimicked that of a living room, a conference room, and anechoic space. Compared to listening through a single microphone, the two-microphone beamformer reduced the target-to-interference ratio required to achieve 50 percent keyword intelligibility by 30, 14, and 0 dB in the anechoic, living-room, and conference-room conditions, respectively. The corresponding improvements over binaural listening (one microphone to each ear) were 24, 9, and 0 dB. Further tests in the living-room environment using the same beam-forming system but with filter impulse responses shortened by a factor of four (which would decrease the adaptation time by a factor of four) decreased the improvement by 5 dB. These results are sufficiently encouraging to warrant further tests involving more realistic reverberant conditions, multiple sources of interference, and time-varying acoustic environments.

Acoustics

Speaking clearly for the hard of hearing. II: Acoustic characteristics of clear and conversational speech.

The first paper of this series (Picheny, Durlach, & Braida, 1985) presented evidence that there are substantial intelligibility differences for hearing-impaired listeners between nonsense sentences spoken in a conversational manner and spoken with the effort to produce clear speech. In this paper, we report the results of acoustic analyses performed on the conversational and clear speech. Among these results are the following. First, speaking rate decreases substantially in clear speech. This decrease is achieved both by inserting pauses between words and by lengthening the durations of individual speech sounds. Second, there are differences between the two speaking modes in the numbers and types of phonological phenomena observed. In conversational speech, vowels are modified or reduced, and word-final stop bursts are often not released. In clear speech, vowels are modified to a lesser extent, and stop bursts, as well as essentially all word-final consonants, are released. Third, the RMS intensities for obstruent sounds, particularly stop consonants, is greater in clear speech than in conversational speech. Finally, changes in the long-term spectrum are small. Thus, speaking clearly cannot be regarded as equivalent to the application of high-frequency emphasis.

Hearing Loss, Sensorineural