PubMed HealthSearch

Biomedical subjects

S A Zahorian

Publications and source records attributed to S A Zahorian.

2 recordsLinked to original sources

Speaker normalization of static and dynamic vowel spectral features.

Two methods are described for speaker normalizing vowel spectral features: one is a multivariable linear transformation of the features and the other is a polynomial warping of the frequency scale. Both normalization algorithms minimize the mean-square error between the transformed data of each speaker and vowel target values obtained from a "typical speaker." These normalization techniques were evaluated both for formants and a form of cepstral coefficients (DCTCs) as spectral parameters, for both static and dynamic features, and with and without fundamental frequency (F0) as an additional feature. The normalizations were tested with a series of automatic classification experiments for vowels. For all conditions, automatic vowel classification rates increased for speaker-normalized data compared to rates obtained for nonnormalized parameters. Typical classification rates for vowel test data for nonnormalized and normalized features respectively are as follows: static formants--69%/79%; formant trajectories--76%/84%; static DCTCs 75%/84%; DCTC trajectories--84%/91%. The linear transformation methods increased the classification rates slightly more than the polynomial frequency warping. The addition of F0 improved the automatic recognition results for nonnormalized vowel spectral features as much as 5.8%. However, the addition of F0 to speaker-normalized spectral features resulted in much smaller increases in automatic recognition rates.

Adult

Vibrotactile frequency for encoding a speech parameter.

Frequency of vibration has not been widely used as a parameter for encoding speech-derived information on the skin. Where it has been used, the frequencies employed have not necessarily been compatible with the capabilities of the tactile channel, and no determination was made of the information transmitted by the frequency variable, as differentiated from other parameters used simultaneously, such as duration, amplitude, and location. However, several investigators have shown that difference limens for vibration frequency may be small enough to make stimulus frequency useful in encoding a speech-derived parameter such as the fundamental frequency of voiced speech. In the studies reported here, measurements have been made of the frequency discrimination ability of the volar forearm, using both sinusoidal and pulse waveforms. Stimulus configurations included the constant-frequency vibrations used by other laboratories as well as frequency-modulated (warbled) stimulus patterns. The frequency of a warbled stimulus was designed to have temporal variations analogous to those found in speech. The results suggest that it may be profitable to display the fundamental frequency of voiced speech on the skin as vibratory frequency, thought it might be desirable to recode fundamental frequency into a frequency range more closely matched to the skin's capability.

Acoustics