PubMed Health⌕ Search

Biomedical subjects

T Baer

Publications and source records attributed to T Baer.

35 records · Page 2Linked to original sources

Spectral contrast enhancement of speech in noise for listeners with sensorineural hearing impairment: effects on intelligibility, quality, and response times.

This paper describes a series of experiments evaluating the effects of digital processing of speech in noise so as to enhance spectral contrast, using subjects with cochlear hearing loss. The enhancement was carried out on a frequency scale related to the equivalent rectangular bandwidths (ERBs) of auditory filters in normally hearing subjects. The aim was to enhance major spectral prominences without enhancing fine-grain spectral features that would not be resolved by a normal ear. In experiment 1, the amount of enhancement and the bandwidth (in ERBs) of the enhancement processing were systematically varied. Large amounts of enhancement produced decreases in the intelligibility of speech in noise. Performance for moderate degrees of enhancement was generally similar to that for the control conditions, possibly because subjects did not have sufficient experience with the processed speech. In experiment 2, subjects judged the relative quality and intelligibility of speech in noise processed using a subset of the conditions of experiment 1. Generally, processing with a moderate degree of enhancement was preferred over the control condition, for both quality and intelligibility. Subjects varied in their preferences for high degrees of enhancement. Experiment 3 used a modified processing algorithm, with a moderate degree of spectral enhancement, and examined the effects of combining the enhancement with dynamic range compression. The intelligibility of speech in noise improved with practice, and, after a small amount of practice, scores for the condition combining enhancement with a moderate degree of compression were found to be significantly higher than for the control condition. Experiment 4 used a subset of conditions from experiment 3, but performance was assessed using a sentence verification test that measured both intelligibility and response times. Scores on both measures were improved by spectral enhancement, and improved still more by enhancement combined with compression. The effects were statistically more robust for the response times. When expressed as equivalent changes in speech-to-noise ratio, the improvements were about twice as large for the response times as for the intelligibility scores. The overall effect of spectral enhancement combined with compression was equivalent to an improvement of speech-to-noise ratio by 4.2 dB.

Adult↗

Analysis of vocal tract shape and dimensions using magnetic resonance imaging: vowels.

Magnetic resonance imaging (MRI) techniques were used to gather basic data to apply in computational models of speech articulation. Two experiments were performed. In experiment 1, voice recordings from two male subjects were obtained simultaneously with axial, coronal, or midsagittal MR images of their vocal tracts while they produced the four point vowels. Area functions describing the individual tract shapes were obtained by measurements performed on the MR images. Digital filters derived from these functions were then used to resynthesize the vowel sounds which were compared, both perceptually and acoustically, with the subjects' original recordings. In experiment 2, axial images of the pharyngeal cavity were collected during the production of an ensemble of nine vowels. Plots of cross-sectional area versus the midsagittal width of the tract at different locations within the pharynx and for different vowel productions were used to derive a functional relationship between the two variables. Data from experiment 1 relating midsagittal width to cross-sectional area within the oral cavity were also examined.

Adult↗

The cricothyroid muscle in voicing control.

Initiation and maintenance of vibrations of the vocal folds require suitable conditions of adduction, longitudinal tension, and transglottal airflow. Thus manipulation of adduction/abduction, stiffening/slackening, or degree of transglottal flow may, in principle, be used to determine the voicing status of a speech segment. This study explores the control of voicing and voicelessness in speech with particular reference to the role of changes in the longitudinal tension of the vocal folds, as indicated by cricothyroid (CT) muscle activity. Electromyographic recordings were made from the CT muscle in two speakers of American English and one speaker of Dutch. The linguistic material consisted of reiterant speech made up of CV syllables where the consonants were voiced and voiceless stops, fricatives, and affricates. Comparison of CT activity associated with the voiced and voiceless consonants indicated a higher level for the voiceless consonants than for their voiced cognates. Measurements of the fundamental frequency (F0) at the beginning of a vowel following the consonant show the common pattern of higher F0 after voiceless consonants. For one subject, there was no difference in cricothyroid activity for voiced and voiceless affricates; in this case, the consonant-induced variations in the F0 of the following vowel were also less robust. Consideration of timing relationships between the EMG curves for voiced and voiceless consonants suggests that the differences most likely reflect control of vocal-fold tension for maintenance or suppression of phonatory vibrations. The same mechanism also seems to contribute to the well-known difference in F0 at the beginning of vowels following voiced and voiceless consonants.

Adult↗

A pitch-synchronous analysis of hoarseness in running speech.

A method of pitch-synchronous acoustic analysis of hoarseness requiring a voice sample of only four fundamental periods is presented. This method calculates a noise-to-signal (N/S) ratio, which indicates the depth of valleys between harmonic peaks in the power spectrum. The spectrum is calculated pitch synchronously from a Fourier transform of the signal, windowed through a continuously variable Hanning window spanning exactly four fundamental periods. A two-stage procedure is used to determine the exact duration of the four fundamental periods. An initial estimate is obtained using autocorrelation in the time domain. A more precise estimate is obtained in the frequency domain by minimizing the errors between the preliminary calculated power spectrum and the predicted spectrum spread of a windowed harmonic signal. Analysis of synthesized voices showed that the N/S ratio is sensitive to additive noise, jitter, and shimmer, and is insensitive to slow (8 Hz) modulation in fundamental frequency and amplitude. An analysis of pre- and postoperative voices of six patients with benign laryngeal disease showed that the N/S ratio for vowel /u/ in running speech consistently improved after surgery for all subjects, in agreement with their successful therapeutic results.

Hoarseness↗

Application of MRI to the analysis of speech production.

Computer models of the process of speech articulation require a detailed knowledge of the vocal tract configurations employed in speech and the application of acoustic theory to calculate the sound waveform. Almost all currently available data on vocal tract dimensions come from x-ray films and are severely limited in quantity and coherence due to restrictions on radiation dosage and intersubject differences. We are using MRI techniques to obtain the pharyngeal dimensions of speakers producing sustained vowels. The fact that MRI does not employ ionizing radiation provides speech research with the opportunity to obtain comprehensive bodies of much-needed data on the articulatory characteristics of single subjects.

Fourier Analysis↗

Frequency and amplitude perturbation analysis of electroglottograph during sustained phonation.

Electroglottography (EGG) was used to monitor vocal fold vibration patterns in normal subjects and patients with various laryngeal disorders. In order to evaluate the regularity of vocal fold vibration, frequency and amplitude perturbation of EGG waves during sustained phonation were measured with a laboratory computer. The data were compared to the degree of hoarseness evaluated by auditory perception and by sound spectrographic analysis. Frequency and amplitude perturbation measures showed some overlap between normal and pathological groups. However, there was a close relation between perturbation analysis of EGG waves and degree of hoarseness (Spearman's rank correlation coefficient rs = 0.73, p less than 0.0005). Amplitude perturbation was found to be a more sensitive measure of the irregularity of vocal fold vibration than frequency perturbation.

Adolescent↗

Onset of voicing in stuttered and fluent utterances.

Electroglottographic (EGG) and acoustic waveforms of the first few glottal pulses of voicing were monitored and voice onset time (VOT) measured during an adaptation task performed by stutterers and controls. The fluent utterances of stutterers resembled those of control subjects. After dysfluencies, however, the EGG signal increased gradually, lending physiological support to the technique of "easy onset" of voicing. EGG waveforms also served to help differentiate mild from severe stutterers. Idiosyncratic ritualized laryngeal behavior, sometimes including physiological tremor, was evident in the EGG record.

Adult↗

Laryngeal vibrations: a comparison between high-speed filming and glottographic techniques.

This study was designed to compare information on laryngeal vibrations obtained by high-speed filming, photoglottography (PGG), and electroglottography (ECG). Simultaneous glottographic signals and high-speed films were obtained from two subjects producing steady phonation. Measurements of glottal width were made at three points along the glottis in the anterior--posterior dimension and aligned with the other records. Results indicate that PGG and film measurements give essentially the same information for peak glottal opening and glottal closure. The EGG signal appears to reliably indicate vocal-fold contact. Together, PGG and EGG may provide much of the information obtained from high-speed filming as well as potentially detect horizontal phase differences during opening and closing.

Electrodiagnosis↗

Some effects of speaking rate on the production of /b/ and /w/.

One of the acoustic properties distinguishing the syllable-initial stop consonant /b/ from the semivowel /w/ is the duration of the initial formant transitions; syllables beginning with /b/ have shorter transitions than those beginning with /w/. This experiment investigated the way in which the transition durations of /b/ and /w/ change as a function of speaking rate by examining tokens of /ba/ and /wa/ produced by four male speakers. At any given speaking rate the /wa/ transitions were, on average, longer than the /ba/ transitions, although pooled across rates, the distributions of transition duration for /ba/ and /wa/ were overlapping. In addition, the magnitude of the difference between average /ba/ and /wa/ transition durations increased with decreases in speaking rate. This is because as rate of speech decreased so that syllable duration increased, there was little change in the initial transition duration of /ba/, but a considerable increase in the initial transition duration of /wa/. Given the overall pattern of results, the transition duration that could optimally distinguish /ba/ from /wa/ was not constant, but increased with syllable duration. This is in accord with Miller and Liberman's (1979) finding that when listeners identify /ba/ and /wa/ on the basis of transition duration, they do so in relation to the duration of the syllable.

Humans↗

Harmonics-to-noise ratio as an index of the degree of hoarseness.

Degree of hoarseness can be evaluated by judging the extent to which noise replaces the harmonic structure in the spectrogram of a sustained vowel. However, this visual method is subjective. The present study was undertaken to develop the harmonics-to-noise (H/N) ratio as an objective and quantitative evaluation of the degree of hoarseness. The computation is conceptually straightforward; 50 consecutive pitch periods of a sustained vowel /a/ are averaged; H is the energy of the averaged waveform, while N is the mean energy of the differences between the individual periods and averaged waveform. Recordings of 42 normal voices and 41 samples with varying degrees of hoarseness were analyzed. Two experts rated the spectrogram of each voice sample, based on the amount of noise relative to that of the harmonic component. The results showed a highly significant agreement (the rank correlation coefficient = 0.849) between H/N calculations and the subjective evaluations of the spectrograms. The H/N ratio also proved useful in quantitatively assessing the results of treatment for hoarseness.

Adult↗

Scaling of glottal opening.

Laryngeal control occurs mainly along two dimensions. One involves the longitudinal tension of the vocal folds and is used for control of fundamental frequency. The other involves abduction/adduction of the folds. This dimension is used in the vegetative functions of the larynx and in its phonetic function to control voicing and aspiration as well as voice quality. Although fine adjustments in timing of abduction/adduction gestures relative to supralaryngeal events produce contrasts of aspiration in obstruents, variations in the size of these gestures appear to be less significant and less finely controlled. The present experiment explores the control of laryngeal abduction/adduction by examining to what extent speakers can control size of glottal aperture under different conditions, with and without suitable feedback. The results suggest that voluntary control of the size of glottal opening is rather poor, and that subjects are unable to make very fine-graded adjustments along this dimension. Voluntary control of glottal opening in isolation is limited, perhaps because normal activities seldom require separate control of this variable. Instead, control of glottal opening is tightly couples to such other activities as respiration, swallowing and speech articulation. Even in the context of sound production, glottal aperture is poorly controlled, perhaps because the precise degree of opening (as opposed to precise timing of opening and closing) has comparatively little practical significance over a wide range of openings.

Glottis↗

A stereo-fiberscope with a magnetic interlens bridge for laryngeal observation.

A stereoscopic method of observation of the larynx and the pharynx during speech utterances has been devised, making use of fiber-optic cables and a magnetic bridge. The cables are inserted via the subject's nostrils. The bridge makes the two objective lenses at the tips of the cables abut within the pharynx near the uvula, and the two images viewed through the separate lenses at the prescribed mutual distance are recorded on each frame of a 16-mm film side by side for computer processing of the three-dimensional data.

Fiber Optic Technology↗

Reflex activation of laryngeal muscles by sudden induced subglottal pressure changes.

In measuring the effect of subglottal pressure changes on fundamental frequency (Fo) of phonation, the effects of changing laryngeal muscle activity must be eliminated. Several investigators have used a strategy in which pulsatile increases of subglottal pressure are induced by pushing on the chest or abdomen of a phonating subject. Fundamental frequency is then correlated with subglottal pressure changes during an interval before laryngeal response is assumed to occur. The present study was undertaken to repeat such an experiment while monitoring electromyographic (EMG) activity of some laryngeal muscles, to discover empirically the latency of the laryngeal response. The results showed a consistent response to each push, with a latency of about 30 ms. Despite this response, analyses of fundamental frequency versus subglottal pressure changes during the interval of constant EMG activity were in general agreement with previously published values. With respect to the nature of the electromyographic response itself, its timing was found to be within the range of latencies appropriate for peripheral feedback, and was also similar to that for an acoustically--or tactually--elicited startle reflex.

Electromyography↗

Voice analysis of the partially ablated larynx. A preliminary report.

This study attempts to obtain a data base of objective formation on the phonatory characteristics of the partially ablated larynx. Twenty patients who had previously undergone partial laryngectomy with glottic reconstruction underwent videolaryngoscopy. The visualizations obtained revealed that the mechanism of voice production was due in part to sphincterization and compensatory hypertrophy of glottic and supraglottic remnants. Aerodynamic and phonatory function tests together with acoustical and perceived voice quality analyses of these partially ablated larynges tend to corroborate the videotape impressions in many instances. However, data accumulated thus far only reveal trends that cannot yet be subjected to definitive interpretations. With the incorporation of other methods of evaluation, augmented by the inclusion of more patient material, it is hoped that the information obtained can be used to improve reconstructive techniques, monitor surgical results, and enhance methods of voice rehabilitation in these patients.

Aged↗

The emphatic and pharyngeal sounds in Hebrew and in Arabic.

This study addresses physiological, acoustic, and linguistic issues in the production of the emphatic sounds [in text] and the pharyngeal sounds [in text]. Approximately 300 minutes of video recordings were obtained from nine Hebrew and Arabic speakers, using a fiberscope positioned in the upper pharynx and simultaneous audio recording through an external microphone. We also studied a cineradiographic film of three Arabic speakers. Results clearly show that all the emphatic sounds, when pronounced as such, share pharyngealization as a secondary articulation. A constriction is formed between the pharyngeal walls and the tip of the epiglottis, which tilts backwards. To a lesser degree, the lower part of the root of the tongue is also retracted. The data show that all the emphatic and pharyngeal sounds we studied are made with qualitatively the same pharyngeal constriction. However, the pharyngeal constriction is more extreme and less variable for the pharyngeal sounds, where it is the primary articulation, than for the emphatic sounds, where it is a secondary articulation. Because the same sort of pharyngealization is seen for all the emphatics, we use a common notational symbol, [in text], for all of them, including [in text] in place of /q/. We note that where pharyngeals and pharyngealized sounds were realized, the Hebrew and Arabic speakers produced them in essentially the same way.

Humans↗