PubMed Health⌕ Search

Biomedical subjects

Q Summerfield

Publications and source records attributed to Q Summerfield.

At least 19 recordsLinked to original sources

Responses of auditory-nerve fibers to stimuli producing psychophysical enhancement.

A form of auditory "enhancement" can be demonstrated by omitting a component from a harmonic series for a few hundred milliseconds and then replacing it: the replaced component stands out perceptually. Psychophysical experiments have shown that components generate more forward masking when enhanced than when present but not enhanced. This result has been interpreted as demonstrating that enhancement involves an increase in gain in the frequency region of the replaced component. The present experiments sought physiological evidence of enhancement in the responses of auditory-nerve fibers in the guinea pig. In one condition a 200-Hz harmonic series lacking components near 2 kHz preceded another series with the 2-kHz component present (the "test" series). In this condition the mean discharge rate to the 2-kHz component was larger than the adapted responses to the other components of the test. In a second condition the test series was preceded by silence. In both conditions the 2-kHz component caused the same increase in firing rate. The average discharge rate sychronized to the 2-kHz component was also the same in both conditions. However, the proportion of the total discharge rate which was locked to 2 kHz was larger when the test followed the harmonic series than when it followed silence. Thus the contrast, in terms of both mean and synchronized rates, between the responses at 2 kHz and those at other frequencies, was increased when the test was preceded by the harmonic series. However, there was no evidence of an increase in gain (i.e., absolutely larger responses) in the 2-kHz region. It seems likely therefore that the mechanisms responsible for this aspect of auditory enhancement are located more central than the auditory nerve.

Acoustic Stimulation↗

The role of frequency modulation in the perceptual segregation of concurrent vowels.

Two experiments investigated the effect of frequency modulation on the identification of vowel sounds presented concurrently with interfering vowels. In experiment 1, identification thresholds were measured for each of five target vowels, masked, in each trial, by one of ten masking vowels. Both target and masking vowels were synthesized using harmonically spaced frequency components. Inharmonic spacing was used in order to prevent powerful grouping processes which exploit fundamental frequency from dominating the results. The target vowels were synthesized with sinusoidal frequency modulation on each frequency component which was either coherent (same phase) or incoherent (random phases). The masking vowels were synthesized with components which were either modulated in the same way as the target vowel or were unmodulated. Identification thresholds were lower when the masking vowel had no modulation. The effect occurred for both coherent and incoherent frequency modulation, indicating that it is mediated by the movement of each component independently, rather than by grouping of coherently modulated components. This result is consistent in some respects with judgments of the prominence of competing vowels [S. E. McAdams, J. Acoust. Soc. Am. 85, 2148-2159 (1989)], which show that modulated vowels are more prominent than unmodulated vowels regardless of the type of modulation applied to the competing vowels. Experiment 2 used a paradigm similar to that developed by McAdams, in order to compare more directly the effect of FM on vowel identification and vowel prominence. On each trial, three vowels were presented concurrently. Either none, one, or two of the vowels were modulated throughout, while modulation was applied to another vowel (the target) halfway through the stimulus. The vowels were either harmonic (with different fundamental frequencies) and coherently modulated or inharmonic and incoherently modulated. Accuracy of identification of the target vowel was not significantly different in the harmonic/coherent and inharmonic/incoherent conditions and declined, in each case, as the number of modulated background vowels increased. Overall, the results of experiments 1 and 2, and of McAdams' prominence judgment data, suggest that there is an auditory mechanism for detecting frequency modulation which can alert the listener to the presence of frequency modulated sounds, but which is insensitive to across-frequency differences in the pattern of that modulation.

Auditory Threshold↗

Perceptual separation of concurrent speech sounds: absence of across-frequency grouping by common interaural delay.

Three experiments and a computational model explored the role of within-channel and across-channel processes in the perceptual separation of competing, complex, broadband sounds which differed in their interaural phase spectra. In each experiment, two competing vowels, whose first and second formants were represented by two discrete bands of noise, were presented concurrently, for identification. Experiments 1 and 2 showed that listeners were able to identify the vowels accurately when each was presented to a different ear, but were unable to identify the vowels when they were presented with different interaural time delays (ITDs); i.e. listeners could not group the noisebands in different frequency regions with the same ITD and thereby separate them from bands in other frequency regions with a different ITD. Experiment 3 demonstrated that while listeners were unable to exploit a difference in interaural delay between the pairs of noisebands, listeners could identify a vowel defined by interaurally decorrelated noisebands when the other two noisebands were interaurally correlated. A computational model based upon that of Durlach [J. Acoust. Soc. Am. 32, 1075-1076 (1960)] showed that the results of these and other experiments can be interpreted in terms of a within-channel mechanism, which is sensitive to interaural decorrelation. Thus the across-frequency integration which occurs in the lateralization of complex sounds may play little role in segregating concurrent sounds.

Humans↗

The contribution of waveform interactions to the perception of concurrent vowels.

Models of the auditory and phonetic analysis of speech must account for the ability of listeners to extract information from speech when competing voices are present. When two synthetic vowels are presented simultaneously and monaurally, listeners can exploit cues provided by a difference in fundamental frequency (F0) between the vowels to help determine their phonemic identities. Three experiments examined the effects of stimulus duration on the perception of such "double vowels." Experiment 1 confirmed earlier findings that a difference in F0 provides a smaller advantage when the duration of the stimulus is brief (50 ms rather than 200 ms). With brief stimuli, there may be insufficient time for attentional mechanisms to switch from the "dominant" member of the pair to the "nondominant" vowel. Alternatively, brief segments may restrict the availability of cues that are distributed over the time course of a longer segment of a double vowel. In experiment 1, listeners did not perform better when the same 50-ms segment was presented four times in succession (with 100-ms silent intervals) rather than only once, suggesting that limits on attention switching do not underlie the duration effect. However, performance improved in some conditions when four successive 50-ms segments were extracted from the 200-ms double vowels and presented in sequence, again with 100-ms silent intervals. Similar improvements were observed in experiment 2 between performance with the first 50-ms segment and one or more of the other three segments when the segments were presented individually. Experiment 3 demonstrated that part of the improvement observed in experiments 1 and 2 could be attributed to waveform interactions that either reinforce or attenuate harmonics that lie near vowel formants. Such interactions were beneficial only when the difference in F0 was small (0.25-1 semitone). These results are compatible with the idea that listeners benefit from small differences in F0 by performing a sequence of analyses of different time segments of a double vowel to determine where the formants of the constituent vowels are best defined.

Adult↗

Clinical evaluation and test-retest reliability of the IHR-McCormick Automated Toy Discrimination Test.

The IHR-McCormick Automated Toy Discrimination Test (ATT) measures the minimum sound level at which a child can identify words presented in quiet in the sound field. This 'word-discrimination threshold' provides a direct measure of the ease with which a child can identify speech and a surrogate measure of auditory sensitivity. This paper describes steps taken to maximize the test-retest reliability of the ATT and to enable it to measure word-discrimination thresholds in noise as well as in quiet. It then describes the results of a clinical evaluation of the ATT in which paediatric audiologists measured word-discrimination thresholds in quiet from 215 successive attendees (in the age range 2 to 13 years) at a paediatric audiology clinic presenting over a 2-month period. When children with atypical cognition or delayed development of language were excluded, 72% of the children provided two word-discrimination thresholds and 83% provided at least one word-discrimination threshold. Children who failed to provide word-discrimination thresholds were generally younger than four years of age. Although a few children who could not perform pure-tone or warble-tone audiometry managed to provide word-discrimination thresholds, most children who could perform the ATT could also perform pure-tone audiometry. The average pure-tone threshold in the better-hearing ear could be predicted from the word-discrimination threshold with a 95% confidence interval of +/- 13 dB. The test-retest reliability of the ATT was measured in two ways. First, to enable comparison with published results, the within-subjects standard deviation of word-discrimination thresholds was calculated. It varied as a function of age and degree of impairment, but was never worse than 3.3 dB. Children of four years of age and older displayed the adult reliability of 2.3 dB. Second, the variability of absolute differences between word-discrimination thresholds was calculated. It was such that a change of 7 dB between two runs of the test (e.g. aided and unaided) would be expected to occur by change less than one time in 20. These results extend previous evaluations of the ATT to a clinically representative population and confirm that word-discrimination thresholds provide a useful complement to warble-tone and pure-tone audiometry.

Adolescent↗

Minimal spectral contrast of formant peaks for vowel recognition as a function of spectral slope.

In four experiments we investigated whether listeners can locate the formants of vowels not only from peaks, but also from spectral "shoulders"--features that give rise to zero crossings in the third, but not the first, differential of the excitation pattern--as hypothesized by Assmann and Summerfield (1989). Stimuli were steady-state approximations to the vowels [a, i, e, u, o] created by summing the first 45 harmonics of a fundamental of 100 Hz. Thirty-nine harmonics had equal amplitudes; the other 6 formed three pairs that were raised in level to define three "formants." An adaptive psychophysical procedure determined the minimal difference in level between the 6 harmonics and the remaining 39 at which the vowels were identifiably different from one another. These thresholds were measured through simulated communication channels, giving overall slopes of the excitation patterns of the five vowels that ranged from -1 dB/erb to + 2 dB/erb. Excitation patterns of the threshold stimuli were computed, and the locations of formants were estimated from zero crossings in the first and third differentials. With the more steeply sloping communication channels, some formants of some vowels were represented as shoulders rather than peaks, confirming the predictions of Assmann and Summerfield's models. We discuss the limitations of the excitation pattern model and the related issue of whether the location of formants can be computed from spectral shoulders in auditory analysis.

Adult↗

Auditory segregation of competing voices: absence of effects of FM or AM coherence.

Four experiments sought evidence that listeners can use coherent changes in the frequency or amplitude of harmonics to segregate concurrent vowels. Segregation was not helped by giving the harmonics of competing vowels different patterns of frequency or amplitude modulation. However, modulating the frequencies of the components of one vowel was beneficial when the other vowel was not modulated, provided that both vowels were composed of components placed randomly in frequency. In addition, staggering the onsets of the two vowels, so that the amplitude of one vowel increased abruptly while the amplitude of the other was stationary, was also beneficial. Thus, the results demonstrate that listeners can group changing harmonics and can segregate them from stationary harmonics, but cannot use coherence of change to separate two sets of changing harmonics.

Acoustic Stimulation↗

Lipreading and audio-visual speech perception.

This paper reviews progress in understanding the psychology of lipreading and audio-visual speech perception. It considers four questions. What distinguishes better from poorer lipreaders? What are the effects of introducing a delay between the acoustical and optical speech signals? What have attempts to produce computer animations of talking faces contributed to our understanding of the visual cues that distinguish consonants and vowels? Finally, how should the process of audio-visual integration in speech perception be described; that is, how are the sights and sounds of talking faces represented at their conflux?

Auditory Perception↗

Effects of signal-to-noise ratio, signal periodicity, and degree of hearing impairment on the performance of voice-separation algorithms.

Procedures for enhancing the intelligibility of a target talker in the presence of a co-channel competing talker were evaluated in tests involving (i) continuously voiced sentences spoken on a monotone, (ii) continuously voiced sentences with time-varying intonation, and (iii) noncontinuously voiced sentences produced with natural intonation. The procedures were based on the methods of harmonic selection and cepstral filtering [R.J. Stubbs and Q. Summerfield, J. Acoust. Soc. Am. 87, 359-372 (1990)]. Target and competing voices were combined at signal-to-noise ratios (SNRs) between -10 dB and +10 dB. Subjects were a group with normal hearing and a heterogeneous group with mild-moderate cochlear hearing impairments. Processing enhanced the target voice over a range of SNRs for each type of sentence and for most listeners. Enhancement was greatest at negative SNRs. Among the impaired listeners, benefit was generally greater for those with milder losses. These results consolidate and extend previous demonstrations that voice-separation algorithms that exploit the harmonic structure of the voiced portions of speech can enhance intelligibility. However, practical application of such algorithms depends on a solution to the problem of tracking the fundamental-frequency contour of one voice in the presence of a competing voice.

Algorithms↗

Perception of concurrent vowels: effects of harmonic misalignment and pitch-period asynchrony.

Three experiments examined the ability of listeners to identify steady-state synthetic vowel-like sounds presented concurrently in pairs to the same ear. Experiment 1 confirmed earlier reports that listeners identify the constituents of such pairs more accurately when they differ in fundamental frequency (f0) by about a half semitone or more, compared to the condition where they have the same f0. When the constituents have different f0's, corresponding harmonics of the two vowels are misaligned in frequency and corresponding pitch periods are asynchronous in time. These differences provide cues that might aid identification. Experiments 2 and 3 determined whether listeners can use these cues, divorced from a difference in f0, to improve their accuracy of identification. Harmonic misalignment was beneficial when the constituents had an f0 of 200 Hz so that the harmonics of each constituent were well separated in frequency. Pitch-period asynchrony was beneficial when the constituents had an f0 of 50 Hz so that the onsets of the pitch periods of each constituent were well separated in time. Neither cue was beneficial when both constituents had an f0 of 100 Hz. It is unlikely, therefore, that either cue contributed to the improvement in performance found in Experiment 1 where the constituents were given different f0's close to 100 Hz. Rather, it is argued that performance improved in Experiment 1 primarily because the two f0's specified two pitches that could be used to segregate the contributions of each vowel in the composite waveform.

Acoustics↗

Algorithms for separating the speech of interfering talkers: evaluations with voiced sentences, and normal-hearing and hearing-impaired listeners.

Two signal-processing algorithms, derived from those described by Stubbs and Summerfield [R.J. Stubbs and Q. Summerfield, J. Acoust. Soc. Am. 84, 1236-1249 (1988)], were used to separate the voiced speech of two talkers speaking simultaneously, at similar intensities, in a single channel. Both algorithms use fundamental frequency (FO) as the basis for segregation. One attenuates the interfering voice by filtering the cepstrum of the signal. The other is a hybrid algorithm that combines cepstral filtering with the technique of harmonic selection [T.W. Parsons, J. Acoust. Soc. Am. 60, 911-918 (1976)]. The algorithms were evaluated and compared in perceptual experiments involving listeners with normal hearing and listeners with cochlear hearing impairments. In experiment 1 the processing was used to separate voiced sentences spoken on a monotone. Both algorithms gave significant increases in intelligibility to both groups of listeners. The improvements were equivalent to an increase of 3-4 dB in the effective signal-to-noise ratio (SNR). In experiment 2 the processing was used to separate voiced sentences spoken with time-varying intonation. For normal-hearing listeners, cepstral filtering gave a significant increase in intelligibility, while the hybrid algorithm gave an increase that was on the margins of significance (p = 0.06). The improvements were equivalent to an increase of 2-3 dB in the effective SNR. For impaired listeners, no intelligibility improvements were demonstrated with intoned sentences. The decrease in performance for intoned material is attributed to limitations of the algorithms when FO is nonstationary.

Algorithms↗

Modeling the perception of concurrent vowels: vowels with different fundamental frequencies.

If two vowels with different fundamental frequencies (fo's) are presented simultaneously and monaurally, listeners often hear two talkers producing different vowels on different pitches. This paper describes the evaluation of four computational models of the auditory and perceptual processes which may underlie this ability. Each model involves four stages: (i) frequency analysis using an "auditory" filter bank, (ii) determination of the pitches present in the stimulus, (iii) segregation of the competing speech sources by grouping energy associated with each pitch to create two derived spectral patterns, and (iv) classification of the derived spectral patterns to predict the probabilities of listeners' vowel-identification responses. The "place" models carry out the operations of pitch determination and spectral segregation by analyzing the distribution of rms levels across the channels of the filter bank. The "place-time" models carry out these operations by analyzing the periodicities in the waveforms in each channel. In their "linear" versions, the place and place-time models operate directly on the waveforms emerging from the filters. In their "nonlinear" versions, analogous operations are applied to the output of an additional stage which applied a compressive nonlinearity to the filtered waveforms. Compared to the other three models, the nonlinear place-time model provides the most accurate estimates of the fo's of paris of concurrent synthetic vowels and comes closest to predicting the identification responses of listeners to such stimuli. Although the model has several limitations, the results are compatible with the idea that a place-time analysis is used to segregate competing sound sources.

Attention↗

A procedure for measuring auditory and audio-visual speech-reception thresholds for sentences in noise: rationale, evaluation, and recommendations for use.

The strategy for measuring speech-reception thresholds for sentences in noise advocated by Plomp and Mimpen (Audiology, 18, 43-52, 1979) was modified to create a reliable test for measuring the difficulty which listeners have in speech reception, both auditorily and audio-visually. The test materials consist of 10 lists of 15 short sentences of homogeneous intelligibility when presented acoustically, and of different, but still homogeneous, intelligibility when presented audio-visually, in white noise. Homogeneity was achieved by applying phonetic and linguistic principles at the stage of compilation, followed by pilot testing and balancing of properties. To run the test, lists are presented at signal-to-noise ratios (SNRs) determined by an up-down psychophysical rule so as to estimate auditory and audio-visual speech-reception thresholds, defined as the SNRs at which the three content words in each sentence are identified correctly on 50% of trials. These thresholds provide measures of a subject's speech-reception abilities. The difference between them provides a measure of the benefit received from vision. It is shown that this measure is closely related to the accuracy with which subjects lip-read words in sentences with no acoustical information. In data from normally hearing adults, the standard deviations (s.d.s) of estimates of auditory speech reception threshold in noise (SRTN), audio-visual SRTN, and visual benefit are 1.2, 2.0, and 2.3 dB, respectively. Graphs are provided with which to estimate the trade-off between reliability and the number of lists presented, and to assess the significance of deviant scores from individual subjects.

Adolescent↗

Strengths and weaknesses of procedures for separating simultaneous voices.

Two signal-processing procedures for separating the continuously-voiced speech of competing talkers are described and evaluated. With competing sentences, each spoken on a monotone, the procedures improved the intelligibility of the target talker both for listeners with normal hearing and for listeners with moderate-to-severe hearing losses of cochlear origin. However, with intoned sentences, benefits were smaller for normal-hearing listeners and were inconsistent for impaired listeners. It is argued that smaller benefits arise with intoned sentences because harmonics of the two voices are blurred together during spectral analysis, limiting the extent to which spectral contrast can be recovered in the processed signal. This is particularly disadvantageous to impaired listeners who have reduced spectro-temporal resolution. This paper discusses other substantial problems to be overcome before the feasibility of the procedures as components of a speech-enhancement system for hearing-impaired listeners could be demonstrated.

Audiometry, Pure-Tone↗

Modeling the perception of concurrent vowels: vowels with the same fundamental frequency.

The ability of listeners to identify pairs of simultaneous synthetic vowels has been investigated in the first of a series of studies on the extraction of phonetic information from multiple-talker waveforms. Both members of the vowel pair had the same onset and offset times and a constant fundamental frequency of 100 Hz. Listeners identified both vowels with an accuracy significantly greater than chance. The pattern of correct responses and confusions was similar for vowels generated by (a) cascade formant synthesis and (b) additive harmonic synthesis that replaced each of the lowest three formants with a single pair of harmonics of equal amplitude. In order to choose an appropriate model for describing listeners' performance, four pattern-matching procedures were evaluated. Each predicted the probability that (i) any individual vowel would be selected as one of the two responses, and (ii) any pair of vowels would be selected. These probabilities were estimated from measures of the similarities of the auditory excitation patterns of the double vowels to those of single-vowel reference patterns. Up to 88% of the variance in individual responses and up to 67% of the variance in pairwise responses could be accounted for by procedures that highlighted spectral peaks and shoulders in the excitation pattern. Procedures that assigned uniform weight to all regions of the excitation pattern gave poorer predictions. These findings support the hypothesis that the auditory system pays particular attention to the frequencies of spectral peaks, and possibly also of shoulders, when identifying vowels. One virtue of this strategy is that the spectral peaks and shoulders can indicate the frequencies of formants when other aspects of spectral shape are obscured by competing sounds.

Humans↗

Auditory enhancement and the perception of concurrent vowels.

Listeners identified both constituents of double vowels created by summing the waveforms of pairs of synthetic vowels with the same duration and fundamental frequency. Accuracy of identification was significantly above chance. Effects of introducing such double vowels by visual or acoustical precursor stimuli were examined. Precursors specified the identity of one of the two constituent vowels. Performance was scored as the accuracy with which the other vowel was identified. Visual precursors were standard English spellings of one member of the vowel pair; acoustical precursors were 1-sec segments of one member of the vowel pair. Neither visual precursors nor contralateral acoustical precursors improved performance over the condition with no precursor. Thus, knowledge of the identity of one of the constituents of a double vowel does not help listeners to identify the other constituent. A significant improvement in performance did occur with ipsilateral acoustical precursors, consistent with earlier demonstrations that frequency components which undergo changes in spectral amplitude achieve enhanced auditory prominence relative to unchanging components. This outcome demonstrates the joint but independent operation of auditory and perceptual processes underlying the ability of listeners to understand speech despite adversely peaked frequency responses in communication channels.

Adult↗

Evaluation of two voice-separation algorithms using normal-hearing and hearing-impaired listeners.

Two signal-processing algorithms, designed to separate the voiced speech of two talkers speaking simultaneously at similar intensities in a single channel, were compared and evaluated. Both algorithms exploit the harmonic structure of voiced speech and require a difference in fundamental frequency (F0) between the voices to operate successfully. One attenuates the interfering voice by filtering the cepstrum of the combined signal. The other uses the method of harmonic selection [T. W. Parsons, J. Acoust. Soc. Am. 60, 911-918 (1976)] to resynthesize the target voice from fragmentary spectral information. Two perceptual evaluations were carried out. One involved the separation of pairs of vowels synthesized on static F0's; the other involved the recovery of consonant-vowel (CV) words masked by a synthesized vowel. Normal-hearing listeners and four listeners with moderate-to-severe, bilateral, symmetrical, sensorineural hearing impairments were tested. All listeners showed increased accuracy of identification when the target voice was enhanced by processing. The vowel-identification data show that intelligibility enhancement is possible over a range of F0 separations between the target and interfering voice. The recovery of CV words demonstrates that the processing is valid not only for spectrally static vowels but also for less intense time-varying voiced consonants. The results for the impaired listeners suggest that the algorithms may be applicable as components of a noise-reduction system in future digital signal-processing hearing aids. The vowel-separation test, and subjective listening, suggest that harmonic selection, which is the more computationally expensive method, produces the more effective voice separation.

Algorithms↗