PubMed Health⌕ Search

Biomedical subjects

K G Munhall

Publications and source records attributed to K G Munhall.

At least 19 recordsLinked to original sources

Remapping auditory-motor representations in voice production.

Evidence regarding visually guided limb movements suggests that the motor system learns and maintains neural maps between motor commands and sensory feedback. Such systems are hypothesized to be used in a feed-forward control strategy that permits precision and stability without the delays of direct feedback control. Human vocalizations involve precise control over vocal and respiratory muscles. However, little is known about the sensorimotor representations underlying speech production. Here, we manipulated the heard fundamental frequency of the voice during speech to demonstrate learning of auditory-motor maps. Mandarin speakers repeatedly produced words with specific pitch patterns (tone categories). On each successive utterance, the frequency of their auditory feedback was increased by 1/100 of a semitone until they heard their feedback one full semitone above their true pitch. Subjects automatically compensated for these changes by lowering their vocal pitch. When feedback was unexpectedly returned to normal, speakers significantly increased the pitch of their productions beyond their initial baseline frequency. This adaptation was found to generalize to the production of another tone category. However, results indicate that a more robust adaptation was produced for the tone that was spoken during feedback alteration. The immediate aftereffects suggest a global remapping of the auditory-motor relationship after an extremely brief training period. However, this learning does not represent a complete transformation of the mapping; rather, it is in part target dependent.

Acoustic Stimulation↗

Visual prosody and speech intelligibility: head movement improves auditory speech perception.

People naturally move their heads when they speak, and our study shows that this rhythmic head motion conveys linguistic information. Three-dimensional head and face motion and the acoustics of a talker producing Japanese sentences were recorded and analyzed. The head movement correlated strongly with the pitch (fundamental frequency) and amplitude of the talker's voice. In a perception study, Japanese subjects viewed realistic talking-head animations based on these movement recordings in a speech-in-noise task. The animations allowed the head motion to be manipulated without changing other characteristics of the visual or acoustic speech. Subjects correctly identified more syllables when natural head motion was present in the animation than when it was eliminated or distorted. These results suggest that nonverbal gestures such as head movements play a more direct role in the perception of speech than previously known.

Adult↗

Spatial frequency requirements for audiovisual speech perception.

Spatial frequency band-pass and low-pass filtered images of a talker were used in an audiovisual speech-in-noise task. Three experiments tested subjects' use of information contained in the different filter bands with center frequencies ranging from 2.7 to 44.1 cycles/face (c/face). Experiment 1 demonstrated that information from a broad range of spatial frequencies enhanced auditory intelligibility. The frequency bands differed in the degree of enhancement, with a peak being observed in a mid-range band (11-c/face center frequency). Experiment 2 showed that this pattern was not influenced by viewing distance and, thus, that the results are best interpreted in object spatial frequency, rather than in retinal coordinates. Experiment 3 showed that low-pass filtered images could produce a performance equivalent to that produced by unfiltered images. These experiments are consistent with the hypothesis that high spatial resolution information is not necessary for audiovisual speech perception and that a limited range of spatial frequency spectrum is sufficient.

Face↗

Impaired processing of prosodic and musical patterns after right hemisphere damage.

The distinction between the processing of musical information and segmental speech information (i.e., consonants and vowels) has been much explored. In contrast, the relationship between the processing of music and prosodic speech information (e.g., intonation) has been largely ignored. We report an assessment of prosodic perception for an amateur musician, KB, who became amusic following a right-hemisphere stroke. Relative to matched controls, KB's segmental speech perception was preserved. However, KB was unable to discriminate pitch or rhythm patterns in linguistic or musical stimuli. He was also impaired on prosodic perception tasks (e.g., discriminating statements from questions). Results are discussed in terms of common neural mechanisms that may underlie the processing of some aspects of both music and speech prosody.

Aged↗

Learning to produce speech with an altered vocal tract: the role of auditory feedback.

Modifying the vocal tract alters a speaker's previously learned acoustic-articulatory relationship. This study investigated the contribution of auditory feedback to the process of adapting to vocal-tract modifications. Subjects said the word /tas/ while wearing a dental prosthesis that extended the length of their maxillary incisor teeth. The prosthesis affected /s/ productions and the subjects were asked to learn to produce "normal" /s/'s. They alternately received normal auditory feedback and noise that masked their natural feedback during productions. Acoustic analysis of the speakers' /s/ productions showed that the distribution of energy across the spectra moved toward that of normal, unperturbed production with increased experience with the prosthesis. However, the acoustic analysis did not show any significant differences in learning dependent on auditory feedback. By contrast, when naive listeners were asked to rate the quality of the speakers' utterances, productions made when auditory feedback was available were evaluated to be closer to the subjects' normal productions than when feedback was masked. The perceptual analysis showed that speakers were able to use auditory information to partially compensate for the vocal-tract modification. Furthermore, utterances produced during the masked conditions also improved over a session, demonstrating that the compensatory articulations were learned and available after auditory feedback was removed.

Adult↗

Gaze behavior in audiovisual speech perception: the influence of ocular fixations on the McGurk effect.

We conducted three experiments in order to examine the influence of gaze behavior and fixation on audiovisual speech perception in a task that required subjects to report the speech sound they perceived during the presentation of congruent and incongruent (McGurk) audiovisual stimuli. Experiment 1 showed that the subjects' natural gaze behavior rarely involved gaze fixations beyond the oral and ocular regions of the talker's face and that these gaze fixations did not predict the likelihood of perceiving the McGurk effect. Experiments 2 and 3 showed that manipulation of the subjects' gaze fixations within the talker's face did not influence audiovisual speech perception substantially and that it was not until the gaze was displaced beyond 10 degrees - 20 degrees from the talker's mouth that the McGurk effect was significantly lessened. Nevertheless, the effect persisted under such eccentric viewing conditions and became negligible only when the subject's gaze was directed 60 degrees eccentrically. These findings demonstrate that the analysis of high spatial frequency information afforded by direct oral foveation is not necessary for the successful processing of visual speech information.

Adolescent↗

Dynamic visual speech perception in a patient with visual form agnosia.

To examine the role of dynamic cues in visual speech perception, a patient with visual form agnosia (DF) was tested with a set of static and dynamic visual displays of three vowels. Five conditions were tested: (1) auditory only which provided only vocal pitch information, (2) dynamic visual only, (3) dynamic audiovisual with vocal pitch information, (4) dynamic audiovisual with full voice information and (5) static visual only images of postures during vowel production. DF showed normal performance in all conditions except the static visual only condition in which she scored at chance. Control subjects scored close to ceiling in this condition. The results suggest that spatiotemporal signatures for objects and events are processed separately from static form cues.

Adult↗

Functional imaging during speech production.

Physiological studies of speech production have demonstrated that even simple articulation involves a range of specialized motor and cognitive processes and the neural mechanisms responsible for speech reflect this complexity. Recently, a number of functional imaging techniques have contributed to our knowledge of the neuroanatomical and neurophysiological correlates of speech production. These new imaging approaches have the advantage of permitting study of large numbers of normal and disordered subjects but they bring with them a host of new methodological concerns. One of the challenges for understanding language production is the recording of articulation itself. The problems associated with measuring the vocal tract and measuring the neural activity during overt speech are reviewed. It is argued that advances in understanding fundamental questions such as what are the planning units of speech, what is the role of feedback during speech and what is the influence of learning, await the development of better methods for assessing task performance.

Brain↗

An inverse dynamics approach to face animation.

Muscle-based models of the human face produce high quality animation but rely on recorded muscle activity signals or synthetic muscle signals that are often derived by trial and error. This paper presents a dynamic inversion of a muscle-based model (Lucero and Munhall, 1999) that permits the animation to be created from kinematic recordings of facial movements. Using a nonlinear optimizer (Powell's algorithm), the inversion produces a muscle activity set for seven muscles in the lower face that minimize the root mean square error between kinematic data recorded with OPTOTRAK and the corresponding nodes of the modeled facial mesh. This inverted muscle activity is then used to animate the facial model. In three tests of the inversion, strong correlations were observed for kinematics produced from synthetic muscle activity, for OPTOTRAK kinematics recorded from a talker for whom the facial model is morphologically adapted and finally for another talker with the model morphology adapted to a different individual. The correspondence between the animation kinematics and the three-dimensional OPTOTRAK data are very good and the animation is of high quality. Because the kinematic to electromyography (EMG) inversion is ill posed, there is no relation between the actual EMG and the inverted EMG. The overall redundancy of the motor system means that many different EMG patterns can produce the same kinematic output.

Adult↗

Perceptual calibration of F0 production: evidence from feedback perturbation.

Hearing one's own speech is important for language learning and maintenance of accurate articulation. For example, people with postlinguistically acquired deafness often show a gradual deterioration of many aspects of speech production. In this manuscript, data are presented that address the role played by acoustic feedback in the control of voice fundamental frequency (F0). Eighteen subjects produced vowels under a control (normal F0 feedback) and two experimental conditions: F0 shifted up and F0 shifted down. In each experimental condition subjects produced vowels during a training period in which their F0 was slowly shifted without their awareness. Following this exposure to transformed F0, their acoustic feedback was returned to normal. Two effects were observed. Subjects compensated for the change in F0 and showed negative aftereffects. When F0 feedback was returned to normal, the subjects modified their produced F0 in the opposite direction to the shift. The results suggest that fundamental frequency is controlled using auditory feedback and with reference to an internal pitch representation. This is consistent with current work on internal models of speech motor control.

Acoustics↗

Anticipatory grip adjustments are observed in both goal-directed movements and movement tics in an individual with Tourette's syndrome.

We examined grip force adjustments during movements of a hand-held object in a young man (BF) with Tourette's syndrome. We directly compared BF's voluntary up and down movements with tics in the same directions. Movement tics were elicited by cueing BF to move either up or down on a GO signal which appeared after a variable delay. During the delay period, we observed frequent tics which were almost always in the cued movement direction. BF's voluntary movements were well coordinated and featured precise and appropriate anticipatory grip force adjustments such that grip force was modulated in phase with movement-induced fluctuations in load. Precise anticipatory grip force adjustments were also observed in all of BF's movement tics. These results support the hypothesis that tics in Tourette's syndrome are purposeful voluntary movements that are well organized and coordinated.

Adaptation, Physiological↗

A model of facial biomechanics for speech production.

Modeling the peripheral speech motor system can advance the understanding of speech motor control and audiovisual speech perception. A 3-D physical model of the human face is presented. The model represents the soft tissue biomechanics with a multilayer deformable mesh. The mesh is controlled by a set of modeled facial muscles which uses a standard Hill-type representation of muscle dynamics. In a test of the model, recorded intramuscular electromyography (EMG) was used to activate the modeled muscles and the kinematics of the mesh was compared with 3-D kinematics recorded with OPTOTRAK. Overall, there was a good match between the recorded data and the model's movements. Animations of the model are provided as MPEG movies.

Biomechanical Phenomena↗

Audiovisual gating and the time course of speech perception.

The time course of audiovisual information in speech perception was examined using a gating paradigm. VCVs that evoked the McGurk effect were gated visually and auditorily. The visual gating yielded a McGurk effect that increased in strength as a linear function of amount of visual stimulus presented. The acoustic gating revealed a more nonlinear function in which the VC information was considerably weaker than the CV portion of the VCV. The results suggest that the flow of cross-modal information in quite complex during audiovisual speech perception.

Humans↗

Temporal characteristics of velopharyngeal function in children.

OBJECTIVE: This investigation was designed to examine the performance of children with normal speech on temporal aspects of aerodynamic tasks related to velopharyngeal closure. DESIGN: The investigation was a descriptive evaluation of variability in aerodynamic features related to velopharyngeal function during multiple repetitions of the word "hamper." SETTING: Children without speech or velopharyngeal difficulties were seen in an experimental laboratory setting for the evaluation procedures. PARTICIPANTS: Twenty-seven subjects were recruited for the experiment. Three subjects were rejected because of behavioral difficulties, and the remaining 24 subjects were subdivided into 4 groups of 6 children (3 males and 3 females) aged 3, 6, 9, and 12 years. The children, who were from local schools and day care centers, volunteered to participate in the experiment. All of the children had age-appropriate speech, language, and hearing abilities, as determined by screening tests administered by one of the examiners (L.T.). MAIN OUTCOME MEASURES: Mean and variability of pressure-flow measures of peak intraoral air pressure and peak nasal airflow and the temporal measures accompanying each air pressure or airflow pulse were evaluated for the age groups of children examined in the experiment. RESULTS: The aerodynamic procedures employed to evaluate velopharyngeal closure during speech were reliable for use with young children. There was a numerical trend toward decreased duration of the temporal parameters with increasing age. Thus, children demonstrated durational values similar to those previously reported for normal-speaking adults. In general, peak oral air pressure and nasal airflow values were like those of previous investigations and demonstrated low variability across all age groups of children tested. CONCLUSIONS: The data from the present investigation provide a preliminary base for comparison of temporal features of velopharyngeal closure for the aerodynamic evaluation of children with impaired velopharyngeal function.

Adult↗

Eye movement of perceivers during audiovisual speech perception.

Perceiver eye movements were recorded during audiovisual presentations of extended monologues. Monologues were presented at different image sizes and with different levels of acoustic masking noise. Two clear targets of gaze fixation were identified, the eyes and the mouth. Regardless of image size, perceivers of both Japanese and English gazed more at the mouth as masking noise levels increased. However, even at the highest noise levels and largest image sizes, subjects gazed at the mouth only about half the time. For the eye target, perceivers typically gazed at one eye more than the other, and the tendency became stronger at higher noise levels. English perceivers displayed more variety of gaze-sequence patterns (e.g., left eye to mouth to left eye to right eye) and persisted in using them at higher noise levels than did Japanese perceivers. No segment-level correlations were found between perceiver eye motions and phoneme identity of the stimuli.

Eye Movements↗

On the registration of time and the patterning of speech movements.

In order to study speech coordination we frequently average kinematic and other physiological signals. The averages are assumed to be more representative of the underlying patterns of production than individual records. In this note we outline different approaches to averaging and present a new nonlinear normalization technique that offers better information than ensemble averaging, linear normalization, or feature alignment methods. We suggest that this technique provides a clear estimation of pattern shape while preserving information on the variation over time.

Humans↗

Functional data analyses of lip motion.

The vocal tract's motion during speech is a complex patterning of the movement of many different articulators according to many different time functions. Understanding this myriad of gestures is important to a number of different disciplines including automatic speech recognition, speech and language pathologies, speech motor control, and experimental phonetics. Central issues are the accurate description of the shape of the vocal tract and determining how each articulator contributes to this shape. A problem facing all of these research areas is how to cope with the multivariate data from speech production experiments. In this paper techniques are described that provide useful tools for describing multivariate functional data such as the measurement of speech movements. The choice of data analysis procedures has been motivated by the need to partition the articulator movement in various ways: end effects separated from shape effects, partitioning of syllable effects, and the splitting of variation within an articulator site from variation from between sites. The techniques of functional data analysis seem admirably suited to the analyses of phenomena such as these. Familiar multivariate procedures such as analysis of variance and principal components analysis have their functional counterparts, and these reveal in a way more suited to the data the important sources of variation in lip motion. Finally, it is found that the analyses of acceleration were especially helpful in suggesting possible control mechanisms. The focus is on using these speech production data to understand the basic principles of coordination. However, it is believed that the tools will have a more general use.

Humans↗

Temporal constraints on the McGurk effect.

Three experiments are reported on the influence of different timing relations on the McGurk effect. In the first experiment, it is shown that strict temporal synchrony between auditory and visual speech stimuli is not required for the McGurk effect. Subjects were strongly influenced by the visual stimuli when the auditory stimuli lagged the visual stimuli by as much as 180 msec. In addition, a stronger McGurk effect was found when the visual and auditory vowels matched. In the second experiment, we paired auditory and visual speech stimuli produced under different speaking conditions (fast, normal, clear). The results showed that the manipulations in both the visual and auditory speaking conditions independently influenced perception. In addition, there was a small but reliable tendency for the better matched stimuli to elicit more McGurk responses than unmatched conditions. In the third experiment, we combined auditory and visual stimuli produced under different speaking conditions (fast, clear) and delayed the acoustics with respect to the visual stimuli. The subjects showed the same pattern of results as in the second experiment. Finally, the delay did not cause different patterns of results for the different audiovisual speaking style combinations. The results suggest that perceivers may be sensitive to the concordance of the time-varying aspects of speech but they do not require temporal coincidence of that information.

Attention↗