PubMed Health⌕ Search

Biomedical subjects

Roy D Patterson

Publications and source records attributed to Roy D Patterson.

At least 19 recordsLinked to original sources

From noise to pitch: transient and sustained responses of the auditory evoked field.

In recent magnetoencephalographic studies, we established a novel component of the auditory evoked field, which is elicited by a transition from noise to pitch in the absence of a change in energy. It is referred to as the 'pitch onset response'. To extend our understanding of pitch-related neural activity, we compared transient and sustained auditory evoked fields in response to a 2000-ms segment of noise and a subsequent 1000-ms segment of regular interval sound (RIS). RIS provokes the same long-term spectral representation in the auditory system as noise, but is distinguished by a definite pitch, the salience of which depends on the degree of temporal regularity. The stimuli were presented at three steps of increasing regularity and two spectral bandwidths. The auditory evoked fields were recorded from both cerebral hemispheres of twelve subjects with a 37-channel magnetoencephalographic system. Both the transient and the sustained components evoked by noise and RIS were sensitive to spectral bandwidth. Moreover, the pitch salience of the RIS systematically affected the pitch onset response, the sustained field, and the off-response. This indicates that the underlying neural generators reflect the emergence, persistence and offset of perceptual attributes derived from the temporal regularity of a sound.

Acoustic Stimulation↗

The effect of temporal context on the sustained pitch response in human auditory cortex.

Recent neuroimaging studies have shown that activity in lateral Heschl's gyrus covaries specifically with the strength of musical pitch. Pitch strength is important for the perceptual distinctiveness of an acoustic event, but in complex auditory scenes, the distinctiveness of an event also depends on its context. In this magnetoencephalography study, we evaluate how temporal context influences the sustained pitch response (SPR) in lateral Heschl's gyrus. In 2 sequences of continuously alternating, periodic target intervals and a more irregular baseline interval, the distinctiveness of the target was decreased in 1 of 2 ways--either by increasing the pitch strength of the baseline or by decreasing the pitch strength of the target. The results show that the amplitude of the SPR increases monotonically with the distinctiveness of the target. Moreover, SPR amplitude is greater for the sequence, where the pitch strength of the target is varied, compared with the condition, where the baseline is varied. Two subsequent experiments show that the amplitude of the SPR increases as duty cycle decreases, in a pitch "strength" contrast and in a pitch "value" contrast. These results indicate that the SPR adapts to recent stimulus history, enhancing the response to rare and brief events.

Acoustic Stimulation↗

Locating the initial stages of speech-sound processing in human temporal cortex.

It is commonly assumed that, in the cochlea and the brainstem, the auditory system processes speech sounds without differentiating them from any other sounds. At some stage, however, it must treat speech sounds and nonspeech sounds differently, since we perceive them as different. The purpose of this study was to delimit the first location in the auditory pathway that makes this distinction using functional MRI, by identifying regions that are differentially sensitive to the internal structure of speech sounds as opposed to closely matched control sounds. We analyzed data from nine right-handed volunteers who were scanned while listening to natural and synthetic vowels, or to nonspeech stimuli matched to the vowel sounds in terms of their long-term energy and both their spectral and temporal profiles. The vowels produced more activation than nonspeech sounds in a bilateral region of the superior temporal sulcus, lateral and inferior to regions of auditory cortex that were activated by both vowels and nonspeech stimuli. The results suggest that the perception of vowel sounds is compatible with a hierarchical model of primate auditory processing in which early cortical stages of processing respond indiscriminately to speech and nonspeech sounds, and only higher regions, beyond anatomically defined auditory cortex, show selectivity for speech sounds.

Adult↗

Comparison of the roex and gammachirp filters as representations of the auditory filter.

Although the rounded-exponential (roex) filter has been successfully used to represent the magnitude response of the auditory filter, recent studies with the roex(p, w, t) filter reveal two serious problems: the fits to notched-noise masking data are somewhat unstable unless the filter is reduced to a physically unrealizable form, and there is no time-domain version of the roex(p, w, t) filter to support modeling of the perception of complex sounds. This paper describes a compressive gammachirp (cGC) filter with the same architecture as the roex(p, w, t) which can be implemented in the time domain. The gain and asymmetry of this parallel cGC filter are shown to be comparable to those of the roex(p, w, t) filter, but the fits to masking data are still somewhat unstable. The roex(p, w, t) and parallel cGC filters were also compared with the cascade cGC filter [Patterson et al., J. Acoust. Soc. Am. 114, 1529-1542 (2003)], which was found to provide an equivalent fit with 25% fewer coefficients. Moreover, the fits were stable. The advantage of the cascade cGC filter appears to derive from its parsimonious representation of the high-frequency side of the filter. It is concluded that cGC filters offer better prospects than roex filters for the representation of the auditory filter.

Acoustics↗

Perception of acoustic scale and size in musical instrument sounds.

There is size information in natural sounds. For example, as humans grow in height, their vocal tracts increase in length, producing a predictable decrease in the formant frequencies of speech sounds. Recent studies have shown that listeners can make fine discriminations about which of two speakers has the longer vocal tract, supporting the view that the auditory system discriminates changes on the acoustic-scale dimension. Listeners can also recognize vowels scaled well beyond the range of vocal tracts normally experienced, indicating that perception is robust to changes in acoustic scale. This paper reports two perceptual experiments designed to extend research on acoustic scale and size perception to the domain of musical sounds: The first study shows that listeners can discriminate the scale of musical instrument sounds reliably, although not quite as well as for voices. The second experiment shows that listeners can recognize the family of an instrument sound which has been modified in pitch and scale beyond the range of normal experience. We conclude that processing of acoustic scale in music perception is very similar to processing of acoustic scale in speech perception.

Acoustic Stimulation↗

The processing and perception of size information in speech sounds.

There is information in speech sounds about the length of the vocal tract; specifically, as a child grows, the resonators in the vocal tract grow and the formant frequencies of the vowels decrease. It has been hypothesized that the auditory system applies a scale transform to all sounds to segregate size information from resonator shape information, and thereby enhance both size perception and speech recognition [Irino and Patterson, Speech Commun. 36, 181-203 (2002)]. This paper describes size discrimination experiments and vowel recognition experiments designed to provide evidence for an auditory scaling mechanism. Vowels were scaled to represent people with vocal tracts much longer and shorter than normal, and with pitches much higher and lower than normal. The results of the discrimination experiments show that listeners can make fine judgments about the relative size of speakers, and they can do so for vowels scaled well beyond the normal range. Similarly, the recognition experiments show good performance for vowels in the normal range, and for vowels scaled well beyond the normal range of experience. Together, the experiments support the hypothesis that the auditory system automatically normalizes for the size information in communication sounds.

Humans↗

The effect of cross-channel synchrony on the perception of temporal regularity.

Temporal models of pitch are based on the assumption that the auditory system measures the time intervals between neural events, and that pitch corresponds to the most common time interval. The current experiments were designed to test whether time intervals are analyzed independently in each peripheral channel, or whether the time-interval analysis in one channel is affected by synchronous activity in other channels. Regular and irregular click trains were filtered into narrow frequency bands to produce target and flanker stimuli. The threshold for discriminating a regular target from an irregular distracter click train was measured in the presence of an irregular masker click train in the target band, as a function of the frequency separation between the target band and a flanker band. The flanker click train was either regular or irregular. The threshold for detecting the regular target was 5-7 dB lower when the flanker was regular. The data indicate that the detection of temporal regularity (and thus, pitch) involves cross-channel processes that can operate over widely separated channels. Model simulations suggest that these cross-channel processes occur after the time-interval extraction stage and that they depend on the similarity, or consistency, of the time-interval patterns in the relevant channels.

Acoustic Stimulation↗

The interaction of glottal-pulse rate and vocal-tract length in judgements of speaker size, sex, and age.

Glottal-pulse rate (GPR) and vocal-tract length (VTL) are related to the size, sex, and age of the speaker but it is not clear how the two factors combine to influence our perception of speaker size, sex, and age. This paper describes experiments designed to measure the effect of the interaction of GPR and VTL upon judgements of speaker size, sex, and age. Vowels were scaled to represent people with a wide range of GPRs and VTLs, including many well beyond the normal range of the population, and listeners were asked to judge the size and sex/age of the speaker. The judgements of speaker size show that VTL has a strong influence upon perceived speaker size. The results for the sex and age categorization (man, woman, boy, or girl) show that, for vowels with GPR and VTL values in the normal range, judgements of speaker sex and age are influenced about equally by GPR and VTL. For vowels with abnormal combinations of low GPRs and short VTLs, the VTL information appears to decide the sex/age judgement.

Acoustic Stimulation↗

Discrimination of speaker size from syllable phrases.

The length of the vocal tract is correlated with speaker size and, so, speech sounds have information about the size of the speaker in a form that is interpretable by the listener. A wide range of different vocal tract lengths exist in the population and humans are able to distinguish speaker size from the speech. Smith et al. [J. Acoust. Soc. Am. 117, 305-318 (2005)] presented vowel sounds to listeners and showed that the ability to discriminate speaker size extends beyond the normal range of speaker sizes which suggests that information about the size and shape of the vocal tract is segregated automatically at an early stage in the processing. This paper reports an extension of the size discrimination research using a much larger set of speech sounds, namely, 180 consonant-vowel and vowel-consonant syllables. Despite the pronounced increase in stimulus variability, there was actually an improvement in discrimination performance over that supported by vowel sounds alone. Performance with vowel-consonant syllables was slightly better than with consonant-vowel syllables. These results support the hypothesis that information about the length of the vocal tract is segregated at an early stage in auditory processing.

Acoustic Stimulation↗

Temporal dynamics of pitch in human auditory cortex.

Recent functional imaging studies have shown that sounds with temporal pitch produce selective activation in anterolateral Heschl's gyrus. This paper reports a magnetoencephalographic (MEG) study of the temporal dynamics of this activation. The cortical response specific to pitch was isolated from the intensity-related response in Planum temporale using a 'continuous stimulation' paradigm in which regular and irregular click trains alternate without interruption. The mean interclick interval (ICI) was 6, 12, 24, or 48 ms; the train length was 720 ms. The auditory sustained field serves as a level-dependent baseline that enhances the signal-to-noise ratio over previous techniques. The onset of pitch was accompanied by a prominent transient field, followed by a strong sustained field, both of which were associated with sources in lateral Heschl's gyrus. The sustained field rose from baseline about 70 ms after the onset of temporal regularity, asymptoted at about 450 ms, and commenced its return to baseline about 70 ms after pitch offset. The peak of the transient field occurred between 130 and 190 ms after regularity onset depending on the ICI. The latencies of the cortical pitch response are substantially longer than might be anticipated from temporal models of pitch perception. This finding suggests that the temporal integration associated with periodicity processing occurs in a subcortical structure, and that the cortical responses reflect subsequent processes involving the measurement of pitch values and changes in pitch.

Auditory Cortex↗

Recovery and refractoriness of auditory evoked fields after gaps in click trains.

When clicks are presented in a train at a rate above approximately 5 Hz, they evoke a sustained field in human auditory cortex that can be recorded by magnetoencephalography. In this study we evaluated how this sustained field continues when a click train is interrupted by a silent gap. The stimuli were click trains with interclick intervals of either 12 or 24 ms, which produce pitches of 83.3 or 41.7 Hz, respectively. The click trains were 996 ms in duration with a gap of 12, 24, 48, 96, or 192 ms beginning 504 ms post-stimulus onset. The sustained field for click trains with short gaps was similar to the one evoked by a continuous click train. Subtraction of the response evoked by a solitary click train of 504 ms enabled estimation of the sustained field in the interval after the gap. The comparison revealed that the sustained field amplitude after the gap was larger than that at the onset of the initial click train in the interval from 150 to 350 ms after onset, and the difference decreased with gap duration. In contrast, the transient P1m was refractory for gaps up to 48 ms, but had nearly recovered its initial amplitude for gaps of 192 ms. We discuss how these results might relate to the perception, i.e. if an interrupted click train is perceived as one continuous sound with a transient gap or as two successive events.

Acoustic Stimulation↗

Mechanisms determining the salience of coloration in echoed sound: influence of interaural time and level differences.

This study investigates whether the salience of the pitch associated with a single reflection of a broadband sound, such as noise, is determined by the monaural information mediated by the stimuli at the two ears, or by the relative locations of the primary sound and the reflection. Pitch strength was measured as a function of the reflection delay and the lateral displacement between the primary sound and the reflection. Thereby, lateral displacement was produced by means of interaural time differences (ITDs) in experiment 1 and interaural level differences (ILDs) in experiment 3. The results from both experiments are in accordance with the assumption that the strength of the pitch associated with a reflection is based on a central average of the internal representations of the stimuli at the two ears. This notion was corroborated by experiment 2, which showed that the results from experiment 1 could be mimicked by simply adding the stimuli from the two ears and presenting the merged stimulus identically to both ears.

Acoustic Stimulation↗

Microsecond temporal resolution in monaural hearing without spectral cues?

The auditory system encodes the timing of peaks in basilar-membrane motion with exquisite precision, and perceptual models of binaural processing indicate that the limit of temporal resolution in humans is as little as 10-20 microseconds. In these binaural studies, pairs of continuous sounds with microsecond differences are presented simultaneously, one sound to each ear. In this paper, a monaural masking experiment is described in which pairs of continuous sounds with microsecond time differences were combined and presented to both ears. The stimuli were matched in terms of the excitation patterns they produced, and a perceptual model of monaural processing indicates that the limit of temporal resolution in this case is similar to that in the binaural system.

Adult↗

Asymmetry of masking between complex tones and noise: partial loudness.

This experiment examined the partial masking of periodic complex tones by a background of noise, and vice versa. The tones had a fundamental frequency (F0) of 62.5 or 250 Hz, and components were added in either cosine phase (CPH) or random phase (RPH). The tones and the noise were bandpass filtered into the same frequency region, from the tenth harmonic up to 5 kHz. The target alone was alternated with the target and the background; for the mixture, the background and target were either gated together, or the background was turned on 400 ms before, and off 200 ms after, the target. Subjects had to adjust the level of either the target alone or the target in the background so as to match the loudness of the target in the two intervals. The overall level of the background was 50 dB SPL, and loudness matches were obtained for several fixed levels of the target alone or in the background. The resulting loudness-matching functions showed clear asymmetry of partial masking. For a given target-to-background ratio, the partial loudness of a complex tone in a noise background was lower than the partial loudness of a noise in a complex tone background. Expressed as the target-to-background ratio required to achieve a given loudness, the asymmetry typically amounted to 12-16 dB. When the F0 of the complex tone was 62.5 Hz, the asymmetry of partial masking was greater for CPH than for RPH. When the F0 was 250 Hz, the asymmetry was greater for RPH than for CPH. Masked thresholds showed the same pattern as for partial masking for both F0's. Onset asynchrony had some effect on the loudness matching data when the target was just above its masked threshold, but did not significantly affect the level at which the target in the background reached its unmasked loudness. The results are interpreted in terms of the temporal structure of the stimuli.

Adult↗

Louder sounds can produce less forward masking: effects of component phase in complex tones.

The influence of the degree of envelope modulation and periodicity on the loudness and effectiveness of sounds as forward maskers was investigated. In the first experiment, listeners matched the loudness of complex tones and noise. The tones had a fundamental frequency (F0) of 62.5 or 250 Hz and were filtered into a frequency range from the 10th harmonic to 5000 Hz. The Gaussian noise was filtered in the same way. The components of the complex tones were added either in cosine phase (CPH), giving a large crest factor, or in random phase (RPH), giving a smaller crest factor. For each F0, subjects matched the loudness between all possible stimulus pairs. Six different levels of the fixed stimulus were used, ranging from about 30 dB SPL to about 80 dB SPL in 10-dB steps. Results showed that, at a given overall level, the CPH and the RPH tones were louder than the noise, and that the CPH tone was louder than the RPH tone. The difference in loudness was larger at medium than at low levels and was only slightly reduced by the addition of a noise intended to mask combination tones. The differences in loudness were slightly smaller for the higher than for the lower F0. In the second experiment, the stimuli with the lower F0s were used as forward maskers of a 20-ms sinusoid, presented at various frequencies within the spectral range of the maskers. Results showed that the CPH tone was the least effective forward masker, even though it was the loudest. The differences in effectiveness as forward maskers depended on masker level and signal frequency; in order to produce equal masking, the level of the CPH tone had to be up to 35 dB above that of the RPH tone and the noise. The implications of these results for models of loudness are discussed and a model is presented based on neural activity patterns in the auditory nerve; this predicts the general pattern of loudness matches. It is suggested that the effects observed in the experiments may have been influenced by two factors: cochlear compression and suppression.

Adult↗

Extending the domain of center frequencies for the compressive gammachirp auditory filter.

The gammatone filter was imported from auditory physiology to provide a time-domain version of the roex auditory filter and enable the development of a realistic auditory filterbank for models of auditory perception [Patterson et al., J. Acoust. Soc. Am. 98, 1890-1894 (1995)]. The gammachirp auditory filter was developed to extend the domain of the gammatone auditory filter and simulate the changes in filter shape that occur with changes in stimulus level. Initially, the gammachirp filter was limited to center frequencies in the 2.0-kHz region where there were sufficient "notched-noise" masking data to define its parameters accurately. Recently, however, the range of the masking data has been extended in two massive studies. This paper reports how a compressive version of the gammachirp auditory filter was fitted to these new data sets to define the filter parameters over the extended frequency range. The results show that the shape of the filter can be specified for the entire domain of the data using just six constants (center frequencies from 0.25 to 6.0 kHz and levels from 30 to 80 dB SPL). The compressive, gammachirp auditory filter also has the advantage of being consistent with physiological studies of cochlear filtering insofar as the compression of the filter is mainly limited to the passband and the form of the chirp in the impulse response is largely independent of level.

Attention↗

Analyzing pitch chroma and pitch height in the human brain.

The perceptual pitch dimensions of chroma and height have distinct representations in the human brain: chroma is represented in cortical areas anterior to primary auditory cortex, whereas height is represented posterior to primary auditory cortex.

Auditory Cortex↗

The processing of temporal pitch and melody information in auditory cortex.

An fMRI experiment was performed to identify the main stages of melody processing in the auditory pathway. Spectrally matched sounds that produce no pitch, fixed pitch, or melody were all found to activate Heschl's gyrus (HG) and planum temporale (PT). Within this region, sounds with pitch produced more activation than those without pitch only in the lateral half of HG. When the pitch was varied to produce a melody, there was activation in regions beyond HG and PT, specifically in the superior temporal gyrus (STG) and planum polare (PP). The results support the view that there is hierarchy of pitch processing in which the center of activity moves anterolaterally away from primary auditory cortex as the processing of melodic sounds proceeds.

Acoustic Stimulation↗