PubMed HealthSearch

PubMed · 6575642

Speech coding: recognizing what we do not hear in speech.

Abstract

Speech is a highly redundant signal. The redundant nature of speech is important for providing reliable communication over air pathways. A large part of this redundancy is useless for speech communication over digital channels. Speech coding aims at minimizing the information rate needed to reproduce a speech signal with specified fidelity. In this paper, we discuss factors that influence the design of efficient speech coders. The encoding and decoding processes invariably introduce error (noise and distortion) in the speech signal. Inability of the human ear to hear certain kinds of distortions in the speech signal plays a crucial role in producing high-quality speech at low bit rates. The physical difference between the waveforms of a given speech signal and its coded replica generally does not tell us much about the subjective quality of the coded signal. A signal-to-noise ratio as small as 10 dB can be tolerated in the coded signal provided the errors are distributed both in time and frequency domains where they are least audible. Recent work on auditory masking has provided us with new insights for optimizing the performance of speech coders. This paper reviews this work and discusses new speech coding methods that attempt to maximize the perceptual similarity between the original speech signal and its coded replica. These new methods make it possible to reproduce speech signals at very low bit rates with little or no audible distortion.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

B S Atal. 1983. Speech coding: recognizing what we do not hear in speech.. https://doi.org/10.1111/j.1749-6632.1983.tb31614.x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

[3-dimensional imaging of temporal bone structures using spiral CT. Initial results in normal temporal bone anatomy].

3D reconstruction of the temporal bone using spiral CT techniques was performed in 51 patients with various otological diseases during routine clinical work evaluation. The 3D display was optimized by a reduced study time and improved detail accuracy by special algorithms. We were able to demonstrate comprehensively in a 3D mode the normal anatomy of the inner ear and adjacent middle ear structures, such as the modiolus of the cochlea, the semicircular canals, the cochlear and vestibular aqueduct and the ossicles. We suggest routine 3D delineation of the substructures of the temporal bone prior to otologic surgery to provide the surgeon with a 3D view of individual anatomy and specific otosurgical sites.

Cochlea

Cochlear implants for congenital deformities.

There have been few accounts of multi-channel cochlear implants in patients with congenital structural deformities of the inner ear which are associated with severe and sometimes progressive deafness. These malformations can now be recognized easily on 2 plane thin section high resolution CT studies which are mandatory for the pre-implantation assessment. However, no attempt seems to have been made to describe which of these malformations would be suitable for an implant or for which would this procedure be contra-indicated. True Mondini deformity of both the cochlea and dilated vestibular aqueduct type would appear suitable for a multi channel implant, but this type of implant should not be used for a primitive otocyst, severe labyrinthine dysplasia or the characteristic X-linked deformity.

Cochlea