PubMed Health⌕ Search

Biomedical subjects

Howard C Nusbaum

Publications and source records attributed to Howard C Nusbaum.

9 recordsLinked to original sources

Hearing lips and seeing voices: how cortical areas supporting speech production mediate audiovisual speech perception.

Observing a speaker's mouth profoundly influences speech perception. For example, listeners perceive an "illusory" "ta" when the video of a face producing /ka/ is dubbed onto an audio /pa/. Here, we show how cortical areas supporting speech production mediate this illusory percept and audiovisual (AV) speech perception more generally. Specifically, cortical activity during AV speech perception occurs in many of the same areas that are active during speech production. We find that different perceptions of the same syllable and the perception of different syllables are associated with different distributions of activity in frontal motor areas involved in speech production. Activity patterns in these frontal motor areas resulting from the illusory "ta" percept are more similar to the activity patterns evoked by AV(/ta/) than they are to patterns evoked by AV(/pa/) or AV(/ka/). In contrast to the activity in frontal motor areas, stimulus-evoked activity for the illusory "ta" in auditory and somatosensory areas and visual areas initially resembles activity evoked by AV(/pa/) and AV(/ka/), respectively. Ultimately, though, activity in these regions comes to resemble activity evoked by AV(/ta/). Together, these results suggest that AV speech elicits in the listener a motor plan for the production of the phoneme that the speaker might have been attempting to produce, and that feedback in the form of efference copy from the motor system ultimately influences the phonetic interpretation.

Auditory Perception↗

The sound of motion in spoken language: visual information conveyed by acoustic properties of speech.

Language is generally viewed as conveying information through symbols whose form is arbitrarily related to their meaning. This arbitrary relation is often assumed to also characterize the mental representations underlying language comprehension. We explore the idea that visuo-spatial information can be analogically conveyed through acoustic properties of speech and that such information is integrated into an analog perceptual representation as a natural part of comprehension. Listeners heard sentences describing objects, spoken at varying speaking rates. After each sentence, participants saw a picture of an object and judged whether it had been mentioned in the sentence. Participants were faster to recognize the object when motion implied by speaking rate matched the motion implied by the picture. Results suggest that visuo-spatial referential information can be analogically conveyed and represented.

Humans↗

Recursive syntactic pattern learning by songbirds.

Humans regularly produce new utterances that are understood by other members of the same language community. Linguistic theories account for this ability through the use of syntactic rules (or generative grammars) that describe the acceptable structure of utterances. The recursive, hierarchical embedding of language units (for example, words or phrases within shorter sentences) that is part of the ability to construct new utterances minimally requires a 'context-free' grammar that is more complex than the 'finite-state' grammars thought sufficient to specify the structure of all non-human communication signals. Recent hypotheses make the central claim that the capacity for syntactic recursion forms the computational core of a uniquely human language faculty. Here we show that European starlings (Sturnus vulgaris) accurately recognize acoustic patterns defined by a recursive, self-embedding, context-free grammar. They are also able to classify new patterns defined by the grammar and reliably exclude agrammatical patterns. Thus, the capacity to classify sequences from recursive, centre-embedded grammars is not uniquely human. This finding opens a new range of complex syntactic processing mechanisms to physiological investigation.

Acoustic Stimulation↗

Repetition suppression for spoken sentences and the effect of task demands.

We examined whether the repeated processing of spoken sentences is accompanied by reduced bold oxygenation level-dependent response (repetition suppression) in regions implicated in sentence comprehension and whether the magnitude of such suppression depends on the task under which the sentences are comprehended or on the complexity of the sentences. We found that sentence repetition was associated with repetition suppression in temporal regions, independent of whether participants judged the sensibility of the statements or listened to the statements passively. In contrast, repetition suppression in inferior frontal regions was found only in the context of the task demanding active judgment. These results suggest that repetition suppression in temporal regions reflects facilitation of sentence comprehension processing per se, whereas in frontal regions it reflects, at least in part, easier execution of specific psycholinguistic judgments.

Adult↗

Listening to talking faces: motor cortical activation during speech perception.

Neurophysiological research suggests that understanding the actions of others harnesses neural circuits that would be used to produce those actions directly. We used fMRI to examine brain areas active during language comprehension in which the speaker was seen and heard while talking (audiovisual) or heard but not seen (audio-alone) or when the speaker was seen talking with the audio track removed (video-alone). We found that audiovisual speech perception activated a network of brain regions that included cortical motor areas involved in planning and executing speech production and areas subserving proprioception related to speech production. These regions included the posterior part of the superior temporal gyrus and sulcus, the pars opercularis, premotor cortex, adjacent primary motor cortex, somatosensory cortex, and the cerebellum. Activity in premotor cortex and posterior superior temporal gyrus and sulcus was modulated by the amount of visually distinguishable phonemes in the stories. None of these regions was activated to the same extent in the audio- or video-alone conditions. These results suggest that integrating observed facial movements into the speech perception process involves a network of multimodal brain regions associated with speech production and that these areas contribute less to speech perception when only auditory signals are present. This distributed network could participate in recognition processing by interpreting visual information about mouth movements as phonetic information based on motor commands that could have generated those movements.

Adult↗

On the neurobiological investigation of language understanding in context.

There are two significant problems in using functional neuroimaging methods to study language. Improving the state of functional brain imaging will depend on understanding how the dependent measure of brain imaging differs from behavioral dependent measures (the "dependent measure problem") and how the activation of the motor system may be confounded with non-motor aspects of processing in certain experimental designs (the "motor output problem"). To address these problems, it may be necessary to shift the focus of language research from the study of linguistic competence to the understanding of language use. This will require investigations of language processing in full multi-modal and environmental context, monitoring of natural behaviors, novel experimental design, and network-based analysis. Such a combined naturalistic approach could lead to tremendous new insights into language and the brain.

Brain Mapping↗

Neural bases of talker normalization.

To recognize phonemes across variation in talkers, listeners can use information about vocal characteristics, a process referred to as "talker normalization." The present study investigates the cortical mechanisms underlying talker normalization using fMRI. Listeners recognized target words presented in either a spoken list produced by a single talker or a mix of different talkers. It was found that both conditions activate an extensive cortical network. However, recognizing words in the mixed-talker condition, relative to the blocked-talker condition, activated middle/superior temporal and superior parietal regions to a greater degree. This temporal-parietal network is possibly associated with selectively attending and processing spectral and spatial acoustic cues required in recognizing speech in a mixed-talker condition.

Acoustic Stimulation↗

Consolidation during sleep of perceptual learning of spoken language.

Memory consolidation resulting from sleep has been seen broadly: in verbal list learning, spatial learning, and skill acquisition in visual and motor tasks. These tasks do not generalize across spatial locations or motor sequences, or to different stimuli in the same location. Although episodic rote learning constitutes a large part of any organism's learning, generalization is a hallmark of adaptive behaviour. In speech, the same phoneme often has different acoustic patterns depending on context. Training on a small set of words improves performance on novel words using the same phonemes but with different acoustic patterns, demonstrating perceptual generalization. Here we show a role of sleep in the consolidation of a naturalistic spoken-language learning task that produces generalization of phonological categories across different acoustic patterns. Recognition performance immediately after training showed a significant improvement that subsequently degraded over the span of a day's retention interval, but completely recovered following sleep. Thus, sleep facilitates the recovery and subsequent retention of material learned opportunistically at any time throughout the day. Performance recovery indicates that representations and mappings associated with generalization are refined and stabilized during sleep.

Humans↗

Selective attention and the acquisition of new phonetic categories.

A class of selective attention models often applied to speech perception is used to study effects of training on the perception of an unfamiliar phonetic contrast. Attention-to-dimension (A2D) models of perceptual learning assume that the dimensions that structure listeners' perceptual space are constant and that learning involves only the reweighting of existing dimensions to emphasize or de-emphasize different sensory dimensions. Multidimensional scaling is used to identify the acoustic-phonetic dimensions listeners use before and after training to recognize the 3 classes of Korean stop consonants. Results suggest that A2D models can account for some observed restructuring of listeners' perceptual space, but listeners also show evidence of directing attention to a previously unattended dimension of phonetic contrast.

Attention↗