PubMed HealthSearch

Biomedical subjects

D G Childers

Publications and source records attributed to D G Childers.

18 recordsLinked to original sources

Detection of laryngeal function using speech and electroglottographic data.

The purpose of this research was to develop quantitative measures for the assessment of laryngeal function using speech and electroglottographic (EGG) data. We developed two procedures for the detection of laryngeal pathology: 1) a spectral distortion measure using pitch synchronous and asynchronous methods with linear predictive coding (LPC) vectors and vector quantization (VQ) and 2) analysis of the EGG signal using time interval and amplitude difference measures. The VQ procedure was conjectured to offer the possibility of circumventing the need to estimate the glottal volume velocity wave-form by inverse filtering techniques. The EGG procedure was to evaluate data that was "nearly" a direct measure of vocal fold vibratory motion and thus was conjectured to offer the potential for providing an excellent assessment of laryngeal function. A threshold based procedure gave 75.9 and 69.0% probability of pathological detection using procedures 1) and 2), respectively, for 29 patients with pathological voices and 52 normal subjects. The false alarm probability was 9.6% for the normal subjects.

Adult

Gender recognition from speech. Part I: Coarse analysis.

The purpose of this research was to investigate the potential effectiveness of digital speech processing and pattern recognition techniques for the automatic recognition of gender from speech segments. In this paper "coarse" acoustic coefficients (autocorrelation, linear prediction, cepstrum, and reflection) were used to form test and reference templates for vowels, voiced fricatives, and unvoiced fricatives. The effects of different distance measures, filter orders, recognition schemes, and vowels and fricatives were comparatively assessed to determine their effectiveness for the task of gender recognition from speech segments. The results showed that most of the acoustic parameters worked well for gender recognition. A within-gender and within-subject averaging technique was important for generating appropriate test and reference templates. The Euclidean distance measure appeared to be the most robust as well as the simplest of the distance measures. The results from this study implied that the gender information is time invariant, phoneme independent, and speaker independent for a given gender. One recognition scheme achieved 100% correct speaker gender classification for a database of 52 talkers (27 male and 25 female). In part II of this paper [D.G. Childers and K. Wu, J. Acoust. Soc. Am. 90, 1841-1856 (1991); hereafter referred to as paper II] the detailed features of ten vowels that appeared responsible for distinguishing a speaker's gender were examined statistically. Included in paper II is a replication of part of the classical study of Peterson and Barney [J. Acoust. Soc. Am. 24, 175-184 (1952)] of vowel characteristics.

Adult

Gender recognition from speech. Part II: Fine analysis.

The purpose of this research was to investigate the potential effectiveness of digital speech processing and pattern recognition techniques for the automatic recognition of gender from speech. In part I Coarse Analysis [K. Wu and D. G. Childers, J. Acoust. Soc. Am. 90, 1828-1840 (1991)] various feature vectors and distance measures were examined to determine their appropriateness for recognizing a speaker's gender from vowels, unvoiced fricatives, and voiced fricatives. One recognition scheme based on feature vectors extracted from vowels achieved 100% correct recognition of the speaker's gender using a database of 52 speakers (27 male and 25 female). In this paper a detailed, fine analysis of the characteristics of vowels is performed, including formant frequencies, bandwidths, and amplitudes, as well as speaker fundamental frequency of voicing. The fine analysis used a pitch synchronous closed-phase analysis technique. Detailed formant features, including frequencies, bandwidths, and amplitudes, were extracted by a closed-phase weighted recursive least-squares method that employed a variable forgetting factor, i.e., WRLS-VFF. The electroglottograph signal was used to locate the closed-phase portion of the speech signal. A two-way statistical analysis of variance (ANOVA) was performed to test the differences between gender features. The relative importance of grouped vowel features was evaluated by a pattern recognition approach. Numerous interesting results were obtained, including the fact that the second formant frequency was a slightly better recognizer of gender than fundamental frequency, giving 98.1% versus 96.2% correct recognition, respectively. The statistical tests indicated that the spectra for female speakers had a steeper slope (or tilt) than that for males. The results suggest that redundant gender information was imbedded in the fundamental frequency and vocal tract resonance characteristics. The feature vectors for female voices were observed to have higher within-group variations than those for male voices. The data in this study were also used to replicate portions of the Peterson and Barney [J. Acoust. Soc. Am. 24, 175-184 (1952)] study of vowels for male and female speakers.

Computer Graphics

Vocal quality factors: analysis, synthesis, and perception.

The purpose of this study was to examine several factors of vocal quality that might be affected by changes in vocal fold vibratory patterns. Four voice types were examined: modal, vocal fry, falsetto, and breathy. Three categories of analysis techniques were developed to extract source-related features from speech and electroglottographic (EGG) signals. Four factors were found to be important for characterizing the glottal excitations for the four voice types: the glottal pulse width, the glottal pulse skewness, the abruptness of glottal closure, and the turbulent noise component. The significance of these factors for voice synthesis was studied and a new voice source model that accounted for certain physiological aspects of vocal fold motion was developed and tested using speech synthesis. Perceptual listening tests were conducted to evaluate the auditory effects of the source model parameters upon synthesized speech. The effects of the spectral slope of the source excitation, the shape of the glottal excitation pulse, and the characteristics of the turbulent noise source were considered. Applications for these research results include synthesis of natural sounding speech, synthesis and modeling of vocal disorders, and the development of speaker independent (or adaptive) speech recognition systems.

Adult

Electroglottography and vocal fold physiology.

The electroglottogram (EGG) is known to be related to vocal fold motion. A major hypothesis undergoing examination in several research centers is that the EGG is related to the area of contact of the vocal folds. This hypothesis is difficult to substantiate with direct measurements using human subjects. However, other supporting evidence can be offered. For this study we made measurements from synchronized ultra high-speed laryngeal films and from EGG waveforms collected from subjects with normal larynges and patients with vocal disorders. We compare certain features of the EGG waveform to (a) the instant of the opening of the glottis, (b) the instant of the closing of the glottis, and (c) the instant of the maximum opening of the glottis. In addition, we compare both the open quotient and the relative average perturbation measured from the glottal area to that estimated from the EGG. All of these comparisons indicate that vocal fold vibratory characteristics are reflected by features of the EGG waveform. This makes the EGG useful for speech analysis and synthesis as well as for modeling laryngeal behavior. The limitations of the EGG are discussed.

Electrodiagnosis

Acoustic correlates of vocal quality.

We have investigated the relationship between various voice qualities and several acoustic measures made from the vowel /i/ phonated by subjects with normal voices and patients with vocal disorders. Among the patients (pathological voices), five qualities were investigated: overall severity, hoarseness, breathiness, roughness, and vocal fry. Six acoustic measures were examined. With one exception, all measures were extracted from the residue signal obtained by inverse filtering the speech signal using the linear predictive coding (LPC) technique. A formal listening test was implemented to rate each pathological voice for each vocal quality. A formal listening test also rated overall excellence of the normal voices. A scale of 1-7 was used. Multiple linear regression analysis between the results of the listening test and the various acoustic measures was used with the prediction sums of squares (PRESS) as the selection criteria. Useful prediction equations of order two or less were obtained relating certain acoustic measures and the ratings of pathological voices for each of the five qualities. The two most useful parameters for predicting vocal quality were the Pitch Amplitude (PA) and the Harmonics-to-Noise Ratio (HNR). No acoustic measure could rank the normal voices.

Adult

Cochannel speech separation.

The multisignal minimum-cross-entropy spectral analysis (multisignal MCESA) is applied to the problem of separating the speech signals of two talkers speaking simultaneously on a single channel, e.g., when two talkers use a single microphone. A new two-stage approach to the problem is proposed in which a spectral separator is followed by a spectral tailoring procedure. The spectral separator produces an initial estimate of the speech spectrum for each talker. Then the spectral tailoring procedure employs the multisignal MCESA technique to adjust the initial spectral estimates to account for the characteristics of the known cochannel composite speech signal. The research emphasis is placed on the implementation and evaluation of the spectral tailoring procedure, i.e., the use of the multisignal MCESA in the proposed scheme. Its usefulness is evaluated and validated by listening tests and by comparing the spectral distortions of the estimated voices before and after the multisignal MCESA processing.

Algorithms

Brain potentials related to seeing one's own name.

Subjects were assigned an assumed name and then shown a series of statements of the form, "My name / is / X", where X was the assumed name, their own first name, or one of a set of other false names. Their task was to respond positively to the "assumed" name and reject as false all other names, including their own. An N380 feature of the averaged task-related brain potentials, considered to be inversely related to the degree of contextual priming, was greatly enhanced for the false names compared to the assumed name. The N380 to one's own name was more similar to that of the false than the assumed name, indicating that the sentence context's priming of various names was under the subjects' attentional control, and that the late negativity could be modulated by this attention. In contrast, a large P510 feature distinguished one's own name from the false name, and this difference was unaffected by practice. Even in cases, then, where the context allows anticipation of one verbal event (here, the assumed name), a highly overlearned and salient stimulus such as one's own name continues to produce a distinctive neural response.

Adolescent

Event-related potentials: a critical review of methods for single-trial detection.

The analysis of ERP data has followed several lines over the last 20 years. The most prevalent method is simply to average ERPs for a given class of stimuli. The ERPs are compared for differences across classes of stimuli. Little other special data processing is used. The ERP comparisons are usually performed using visual examination of the wave-shapes. Sometimes statistics are calculated such as means, variances, and confidence limits. Linear filtering is used to reduce interference. Another approach is to model or analyze the ERP as a sequence of vectors or frames of data samples. These samples may be of the ERP time waveform or they may be of the frequency transform of the ERP waveform. The frames of data vary in length from the entire ERP waveform (500 to 1000 msec) to frames as short as ten sample points (100 msec). Recognition of an event in the ERP is achieved by computing a distance measure between parameter vectors for one class of stimuli and corresponding parameter vectors for another class of stimuli. Recognition is achieved by selecting the ERP with the lowest distance score. This approach is "pattern matching" and relies on two assumptions: adjacent frames of data are uncorrelated, and the variability of the data can be accounted for by the distance measured for all stimuli in the classes presented. Subject variability is generally not accounted for, other than to assume it is the same for all classes of stimuli. The data are clustered into a variety of reference patterns that represent particular manifestations of a particular stimulus. Another approach is "feature-based" recognition. The idea is to identify and automatically extract features of the data that can provide a characterization of stimuli. The features selected may be abstract. They are calculated from the data or transforms of the data.

Biometry

A model for vocal fold vibratory motion, contact area, and the electroglottogram.

The electroglottogram (EGG) has been conjectured to be related to the area of contact between the vocal folds. This hypothesis has been substantiated only partially via direct and indirect observations. In this paper, a simple model of vocal fold vibratory motion is used to estimate the vocal fold contact area as a function of time. This model employs a limited number of vocal fold vibratory features extracted from ultra high-speed laryngeal films. These characteristics include the opening and closing vocal fold angles and the lag (phase difference) between the upper and lower vocal fold margins. The electroglottogram is simulated using the contact area, and the EGG waveforms are compared to measured EGGs for normal male voices producing both modal and pulse register tones. The model also predicts EGG waveforms for vocal fold vibration associated with a nodule or polyp.

Glottis

Brain potentials during sentence verification: automatic aspects of comprehension.

College students learned a set of facts relating fictitious people and their occupations (e.g. 'Matthew is a lawyer'). Event-related brain potentials (ERPs) were recorded while they subsequently viewed a series of such statements presented in segments (e.g. 'Matthew/is a/dentist'). ERPs to occupations completing statements falsely were significantly more negative than those to true statements in an interval 200-420 msec poststimulus (peak N320), whether subjects were required to make a decision about each statement or passively view the presented segments (Experiments 1 and 2). A later ERP positivity was observed during 'response' trials that was of longer latency for false than true completions; but this positive component was greatly attenuated during 'no-response' trials. The enhanced N320 for false completions was not affected by requiring subjects on some trials to respond incorrectly (Experiment 3). It is concluded that attending to a presented word results in an automatic analysis of its meaning in the context of a preceding verbal input, and that ERPs can indicate the nature of the output of that analysis.

Brain

A critical review of electroglottography.

The technique of electroglottography is reviewed from the perspective of a laboratory instrument for assessing laryngeal function, a device to assist speech and speaker recognition, and as a potential diagnostic aid in the clinic. A description of the electronic functioning of the electroglottograph (EGG) is provided. Considerable emphasis is given to contemporary research which has focused on laryngeal assessment using the EGG. Methods for validating and aiding the interpretation or reading of the EGG are discussed, including photoglottography, stroboscopy, ultrahigh-speed laryngeal cinematography, and others. The relationship of the EGG to glottal area and glottal volume velocity estimated by inverse filtering is presented. An elementary model of the EGG is described and used to predict characteristic features of the EGG waveform. Clinical data as well as data obtained from subjects with a normal functioning larynx are analyzed. Applications of the EGG to speech processing are outlined, including real-time detection of voicing, voiced and unvoiced speech segments, and silence intervals. The EGG device has potential for assisting speech and speaker recognition systems in certain applications.

Aged

Laryngeal vibration patterns. Machine-aided measurements from high-speed film.

A report is given on the development of procedures to process laryngeal high-speed films to extract the glottal wave-form and other glottal measurements. The glottal waveforms that are derived by this direct method are being used to determine whether the pathological larynx is manifested in its abnormal vibratory pattern. The glottal waveforms measured by sonic-sensing pen tracing, cursor outlining, a photocell technique, and television camera scanning are presented and compared with the conventional polar planimeter method. To date, the television camera method is the most rapid procedure and can process approximately 400 frames per hour in a semiautomatic mode. This method determines the glottal area, length, and width by a combination of analog-digital circuitry and minicomputer processing. This information is being used to explicate the exact nature of the vibrational patterns produced by the normal and pathological larynges.

Computers