PubMed Health⌕ Search

PubMed · 12199395

Speech synthesis using damped sinusoids.

Abstract

A speech synthesizer was developed that operates by summing exponentially damped sinusoids at frequencies and amplitudes corresponding to peaks derived from the spectrum envelope of the speech signal. The spectrum analysis begins with the calculation of a smoothed Fourier spectrum. A masking threshold is then computed for each frame as the running average of spectral amplitudes over an 800-Hz window. In a rough simulation of lateral suppression, the running average is then subtracted from the smoothed spectrum (with negative spectral values set to zero), producing a masked spectrum. The signal is resynthesized by summing exponentially damped sinusoids at frequencies corresponding to peaks in the masked spectra. If a periodicity measure indicates that a given analysis frame is voiced, the damped sinusoids are pulsed at a rate corresponding to the measured fundamental period. For unvoiced speech, the damped sinusoids are pulsed on and off at random intervals. A perceptual evaluation of speech produced by the damped sinewave synthesizer showed excellent sentence intelligibility, excellent intelligibility for vowels in /hVd/ syllables, and fair intelligibility for consonants in CV nonsense syllables.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

James M Hillenbrand, Robert A Houde. 2002. Speech synthesis using damped sinusoids.. https://doi.org/10.1044/1092-4388(2002%2F051)

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Was blind but now I see.

Explore the source record for details and available documents.

Communication Devices for People with Disabilities↗

Communication with deaf and hard-of-hearing people: a guide for medical education.

Some physicians may be insufficiently prepared to work with the many patients who have hearing loss. People with hearing loss constitute approximately 9% of the U.S. population, and the prevalence is increasing. Patients with hearing loss and their physicians report communication difficulties; physicians also report feeling less comfortable with these patients. Although communication with patients plays a major role in determining diagnoses and management, little attention is given to teaching medical students and residents the skills necessary to facilitate communication when hearing loss is involved. The need for these skills will increase with the expected rise in the number of such patients. The author presents the rationale for including information about hearing loss in curricula on patient-doctor communication, and suggests curricular content, including background regarding hearing loss and techniques that can enhance the physician's ability to listen to (that is, "hear") and learn about the stories of these patients.

Communication Devices for People with Disabilities↗

A quasiarticulatory approach to controlling acoustic source parameters in a Klatt-type formant synthesizer using HLsyn.

The HLsyn speech synthesizer uses models of the vocal tract to map higher-level quasiarticulatory parameters to the acoustic parameters of a Klatt-type formant synthesizer. The benefits of this system are several. In addition to requiring a relatively small number of parameters, the HLsyn model includes constraints on source-filter relations that occur naturally during speech production. Such constraints help to prevent combinations of sources and filter that are impossible to achieve with the human vocal tract. Thus, HLsyn could lead to reductions in the complexity of formant synthesis and result in better quality synthesis. HLsyn can also be a useful tool for speech-science education and speech research. This paper focuses on the generation of acoustic sources in HLsyn. Described in detail are the equations and methods used to estimate Klatt-type source parameters from HLsyn parameters. Several examples illustrating the generation of source parameters for obstruents (voiced and voiceless) and sonorants are provided. Future papers will describe the filtering components of HLsyn.

Communication Devices for People with Disabilities↗