PubMed Health⌕ Search

Biomedical subjects

Martin Cooke

Publications and source records attributed to Martin Cooke.

4 recordsLinked to original sources

A glimpsing model of speech perception in noise.

Do listeners process noisy speech by taking advantage of "glimpses"-spectrotemporal regions in which the target signal is least affected by the background? This study used an automatic speech recognition system, adapted for use with partially specified inputs, to identify consonants in noise. Twelve masking conditions were chosen to create a range of glimpse sizes. Several different glimpsing models were employed, differing in the local signal-to-noise ratio (SNR) used for detection, the minimum glimpse size, and the use of information in the masked regions. Recognition results were compared with behavioral data. A quantitative analysis demonstrated that the proportion of the time-frequency plane glimpsed is a good predictor of intelligibility. Recognition scores in each noise condition confirmed that sufficient information exists in glimpses to support consonant identification. Close fits to listeners' performance were obtained at two local SNR thresholds: one at around 8 dB and another in the range -5 to -2 dB. A transmitted information analysis revealed that cues to voicing are degraded more in the model than in human auditory processing.

Acoustic Stimulation↗

Effect of masker type on native and non-native consonant perception in noise.

Spoken communication in a non-native language is especially difficult in the presence of noise. This study compared English and Spanish listeners' perceptions of English intervocalic consonants as a function of masker type. Three maskers (stationary noise, multitalker babble, and competing speech) provided varying amounts of energetic and informational masking. Competing English and Spanish speech maskers were used to examine the effect of masker language. Non-native performance fell short of that of native listeners in quiet, but a larger performance differential was found for all masking conditions. Both groups performed better in competing speech than in stationary noise, and both suffered most in babble. Since babble is a less effective energetic masker than stationary noise, these results suggest that non-native listeners are more adversely affected by both energetic and informational masking. A strong correlation was found between non-native performance in quiet and degree of deterioration in noise, suggesting that non-native phonetic category learning can be fragile. A small effect of language background was evident: English listeners performed better when the competing speech was Spanish.

Adolescent↗

An audio-visual corpus for speech perception and automatic speech recognition.

An audio-visual corpus has been collected to support the use of common material in speech perception and automatic speech recognition studies. The corpus consists of high-quality audio and video recordings of 1000 sentences spoken by each of 34 talkers. Sentences are simple, syntactically identical phrases such as "place green at B 4 now". Intelligibility tests using the audio signals suggest that the material is easily identifiable in quiet and low levels of stationary noise. The annotated corpus is available on the web for research use.

Acoustic Stimulation↗

Consonant identification in N-talker babble is a nonmonotonic function of N.

Consonant identification rates were measured for vowel-consonant-vowel tokens gated with N-talker babble noise and babble-modulated noise for an extensive range of N, at a fixed signal-to-noise ratio. In the natural babble condition, intelligibility was a nonmonotonic function of N, with a broad performance minimum from N = 6 to N = 128. Identification rates in babble-modulated noise fell gradually with N. The contributions of factors such as energetic masking, linguistic confusion, attentional load, peripheral adaptation, and stationarity to the perception of consonants in N-talker babble are discussed.

Acoustic Stimulation↗