PubMed HealthSearch

Biomedical subjects

D H Whalen

Publications and source records attributed to D H Whalen.

At least 19 recordsLinked to original sources

Predicting midsagittal pharynx shape from tongue position during vowel production.

The shape of the pharynx has a large effect on the acoustics of vowels, but direct measurement of this part of the vocal tract is difficult. The present study examines the efficacy of inferring midsagittal pharynx shape from the position of the tongue, which is much more amenable to measurement. Midsagittal magnetic resonance (MR) images were obtained for multiple repetitions of 11 static English vowels spoken by two subjects (one male and one female). From these, midsagittal widths were measured at approximately 3-mm intervals along the entire vocal tract. A regression analysis was then used to assess whether the pharyngeal widths could be predicted from the locations and width measurements for four positions on the tongue, namely, those likely to be the locations of a receiver coil for an electromagnetometer system. Predictability was quite high throughout the vocal tract (multiple r> 0.9), except for the extreme ends (i.e., larynx and lips) and small decreases for the male subject in the uvula region. The residuals from this analysis showed that the accuracy of predictions was generally quite high, with 89.2% of errors being less than 2 mm. The extremes of the vocal tract, where the resolution of the MRI was poorer, accounted for much of the error. For languages like English, which do not use advanced tongue root (ATR) distinctively, the midsagittal pharynx shape of static vowels can be predicted with high accuracy.

Female

Exploring the relationship of inspiration duration to utterance duration.

Previous work has indicated that there may be a positive relationship between the duration and extent of inspiration and the length of an upcoming utterance. However, none of that work has uniquely implied a role of planning. We attempted to avoid some of the alternative explanations by forcing subjects to utter single sentences ranging in length from 5 to 82 syllables (mean of 27), after inspiring fully and then expiring down to a set level before uttering the sentence. For all 3 subjects, there was a significant positive relationship between utterance length and inspiration duration, regardless of whether inspiration was measured physiologically or acoustically. The 2 subjects with the higher correlations in the articulatory measures also expended air more quickly during the shorter sentences than longer ones, while the other subject had no correlation with exhalation rate. Complexity of the sentence, calculated as the number of clauses in the sentence, did not affect inspiration duration. The individual differences need further investigation, but there is a positive correlation between the duration of the sentence to be said and the inspiration before it when the speaker is required to read sentences while using only one breath.

Humans

Limits on phonetic integration in duplex perception.

The telling fact about duplex perception is that listeners integrate into a unitary phonetic percept signals that are coherent from a phonetic point of view, even though the signals are, on purely auditory grounds, separate sources. Here we explore the limits on the integration of a sinusoidal consonant cue (the F3 transition for [da] vs. [ga]) with the resonances of the remainder of the syllable. Perceiving duplexly, listeners hear the whistle of the sinusoid, but also the [da] and [ga] for which the sinusoid provides the critical information. In the first experiment, phonetic integration was significantly reduced, but not to zero, by a precursor that extended the transition cue forward in time so that it started 50 msec before the cue. The effect was the same above and below the duplexity threshold (the intensity of sinusoid in the combined pattern at which the whistle was just barely audible). In the second experiment, integration was reduced once again by the precursor, and also, but only below the duplexity threshold, by harmonics of the cues that were simultaneous with it. The third experiment showed that the simultaneous harmonics reduced phonetic integration only by serving as distractors while also permitting the conclusion that the precursor produced its effects by making the cue part of a coherent and competing auditory pattern, and so "capturing" it. The fourth experiment supported this interpretation by showing that for some subjects the amount of capture was reduced when the capturing tone was itself captured by being made part of a tonal complex. The results support the assumption that the independent phonetic system will integrate across disparate sources according to the cohesive power of that system as measured against the evidence for separate sources.

Adult

The effects of breath sounds on the perception of synthetic speech.

When preparing to speak, talkers typically take a breath. The perceptual effect of adding naturally produced breath intake sounds to synthetic speech was examined. In experiment 1, subjects were better at transcribing synthesized sentences that were preceded by a breath sound than those that were not, in addition to the improvement due to practice that is typically found with synthetic speech. Experiment 2 found that replacing the breath with the spectrally similar sound of rustling leaves had no effect on the accuracy. Experiment 3 had breaths before randomly selected sentences. Only the practice effect was significant, though there was a tendency for sentences with the breath sounds to be remembered better. In experiment 4, we tested whether the appropriateness of the breath sound to the sentence size (relatively short or long) affected the use of the breath sound. Appropriateness had no effect, perhaps because the range of sentence durations was too small. Experiment 5 replicated experiment 1 but used leaf sounds rather than silence in the nonbreath sentences. The presence of breath was again found to aid recall. Overall, the current results indicate that adding the breath intake sound to synthetic sentences improves listeners' ability to recall those sentences.

Female

Intrinsic F0 of vowels in the babbling of 6-, 9-, and 12-month-old French- and English-learning infants.

In every language so far examined, high vowels such as [i] and [u] tend to have higher fundamental frequencies (F0s) than low vowels such as [a]. This intrinsic F0 effect (IF0) has been found in the speech of children at various stages of development, except in the one previous study of babbling. The present study is based on a larger set of utterances from more subjects (six French- and six English-learning infants), at the ages 6, 9, and 12 months. It is found, instead, that IF0 appears even in babbling. There is no indication in these data of a developmental trend for the effect, and no indication of a difference due to the target language. These results support the claim that IF0 is an automatic consequence of producing vowels.

Child Development

FO gives voicing information even with unambiguous voice onset times.

The voiced/voiceless distinction for English utterance-initial stop consonants is primarily realized as differences in the voice onset time (VOT), which is largely signaled by the time between the stop burst and the onset of voicing. The voicing of stops has also been shown to affect the vowel's FO after release, with voiceless stops being associated with higher FO. When the VOT is ambiguous, these FO "perturbations" have been shown to affect voicing judgments. This is to be expected of what can be considered a redundant feature, that is, that it should carry a distinction in cases where the primary feature is neutralized. However, when the voicing judgments were made as quickly as possible, an inappropriate FO was found to slow response time even for unambiguous VOTs. This was true both of FO contours and level FO differences. These results reinforce the plausibility of tonogenesis, and they add further weight to the claim that listeners make full use of the signal given to them, even when overt labeling would seem to indicate otherwise.

Audiometry

Information for Mandarin tones in the amplitude contour and in brief segments.

While the tones of Mandarin are conveyed mainly by the F0 contour, they also differ consistently in duration and in amplitude contour. The contribution of these factors was examined by using signal-correlated noise stimuli, in which natural speech is manipulated so that it has no F0 or formant structure but retains its original amplitude contour and duration. Tones 2, 3 and 4 were perceptible from just the amplitude contour, even when duration was not also a cue. In two further experiments, the location of the critical information for the tones during the course of the syllable was examined by extracting small segments from each part of the original syllable. Tones 2 and 3 were often confused with each other, and segments which did not have much F0 change were most often heard as Tone 1. There were, though, also cases in which a low, unchanging pitch was heard as Tone 3, indicating a partial effect of register even in Mandarin. F0 was positively correlated with amplitude, even when both were computed on a pitch period basis. Taken together, the results show that Mandarin tones are realized in more than just the F0 pattern, that amplitude contours can be used by listeners as cues for tone identification, and that not every portion of the F0 pattern unambiguously indicates the original tone.

Adult

Intonational differences between the reduplicative babbling of French- and English-learning infants.

The two- and three-syllable reduplicative babbling of five French-learning and five English-learning infants (0;5 to 1;1) was examined in two ways for intonational differences. The first measure was a categorization into one of five categories (RISING, FALLING, RISE-FALL, FALL-RISE, LEVEL) by expert listeners. The second was the fundamental frequency (F0) from the early, middle and late portion of each syllable. Both measures showed significant differences between the two language groups. 65% of the utterances from both groups were classified as either rising of falling. For the French children, these were divided equally into the rising and the falling categories, while 75% of those utterances for the English children were judged to have falling intonation. Proportions of the other three categories were not significantly different by language environment. In both languages, though, three-syllable utterances were more likely to have a complex contour than two-syllable ones. Analysis of the F0 patterns confirmed the perceptual assessment. Several aspects of the target languages help explain these intonational differences in prelinguistic babbling.

Data Interpretation, Statistical

Perception of the English /s/-/integral of/ distinction relies on fricative noises and transitions, not on brief spectral slices.

A series of experiments compared two approaches to fricative identification, spectral template matching and articulatory dynamics. Natural-speech /s/ and /integral of/ noises from fricative-vowel or vowel-fricative syllables were cross spliced so that "hybrid" noises started out as either /s/ or /integral of/ and ended up with the other fricative noise in varying proportions. With both initial and final fricatives, listener judgments most often agreed with the longer part of the noise even when spectral templates would predict the other category. Also, the vocalic formant transitions contributed to the judgment. In another experiment, open transcriptions by four expert listeners similarly showed that all the cues were used; there were also some instances of nonspeech percepts that would be predicted by gestural models. One further experiment had subjects identify two fricatives from hybrid noises between two vocalic segments. When the order of the noises differed from the order of the transitions, the perceived ordering of the fricatives was often the reverse of the order of the noise segments. Taken together with previous results, these experiments indicate that listeners take the whole fricative noise, as well as the transitions, into account in fricative identification.

Adult

Subcategorical phonetic mismatches and lexical access.

The place of phonetic analysis in the perception of words is unclear. While some theories assume fully specified phonemic strings as input, other theories assume that little analysis occurs. An earlier experiment by Streeter and Nigro (1979) produced evidence, based on auditorily presented words with misleading acoustic cues, that lexical decisions were based on mostly unanalyzed patterns, since word judgments were delayed by misleading information whereas nonword judgments were not. The present studies expand that work to a different set of cues, and to cases in which the overriding cue came first. An additional task, auditory naming, was used to examine the effects when the decision stage is less demanding. For the lexical decision task, misleading information slowed the responses, for both words and nonwords. In the auditory naming task, only the slower responses were affected. These results suggest that phonetic conflicts are resolved prior to lexical access.

Adolescent

Gradient effects of fundamental frequency on stop consonant voicing judgments.

The post-stop-release rise or fall of fundamental frequency (F0) is known to affect voicing judgments of syllables with ambiguous voice onset times (VOTs). In 1986, Silverman claimed that the critical factor was not direction of F0 change but rather its direction relative to the intonational contour. He further claimed that only F0s that start above and fall to the contour have an effect proportional to the size of the frequency change; F0s that rise to the contour by different amounts were claimed to be equivalent. In our first experiment, we examined the effect on voicing judgments of five onset F0s preceding a single, flat contour. Only falling F0s were differentiated in the first set of judgments, but after increased exposure to the syllables, even F0s below the contour differentially affected the voicing judgment. In a second experiment, the contour of the final part of the syllable was flat, rising or falling. F0 contour affected the judgments, as did onset F0s, but the two factors did not interact, indicating that the onset values were not being judged by reference to the contours. However, the contour which was predicted to result in more voiceless judgments also ended at a higher F0 in the vowel, and another effect of voicing is that the F0 is higher throughout the vowel after voiceless stops. In a third experiment, F0 contours were created to contrast contour and mean F0. The effect of the F0 during the vocalic segment appeared to be attributable to the average F0 rather than the contour. In all three experiments, the F0 onset values contributed to the voicing judgment whether they were above or below the putative intonation contour. The contribution of the lower F0s, while significant, was not as great as that of the higher F0s, which argues for a noncategorical contribution of intonation.

Cues

The perceptual effects of child-adult differences in fricative-vowel coarticulation.

Earlier work [Nittrouer et al., J. Speech Hear. Res. 32, 120-132 (1989)] demonstrated greater evidence of coarticulation in the fricative-vowel syllables of children than in those of adults when measured by anticipatory vowel effects on the resonant frequency of the fricative back cavity. In the present study, three experiments showed that this increased coarticulation led to improved vowel recognition from the fricative noise alone: Vowel identification by adult listeners was better overall for children's productions and was successful earlier in the fricative noise. This enhanced vowel recognition for children's samples was obtained in spite of the fact that children's and adults' samples were randomized together, therefore indicating that listeners were able to normalize the vowel information within a fricative noise where there often was acoustic evidence of only one formant associated primarily with the vowel. Correct vowel judgments were found to be largely independent of fricative identification. However, when another coarticulatory effect, the lowering of the main spectral prominence of the fricative noise for /u/ versus /i/, was taken into account, vowel judgments were found to interact with fricative identification. The results show that listeners are sensitive to the greater coarticulation in children's fricative-vowel syllables, and that, in some circumstances, they do not need to make a correct identification of the most prominently specified phone in order to make a correct identification of a coarticulated one.

Adult

P-center judgments are generally insensitive to the instructions given.

The perceptual moment of occurrence of a syllable, its P-center, has frequently been examined by instructing subjects to adjust a series of speech sounds until they sounded isochronous. The present two experiments examined the effect of changing the instructions. In addition to the overall isochrony instructions, we asked subjects to align pairs of syllables so that the syllable onsets, vowel onsets, or syllable offsets sounded isochronous. In the first experiment, 3 of 4 subjects showed no difference among the first three instruction sets, and the changes introduced by the fourth went in the wrong direction. All subjects found it impossible to make alignments with respect to offsets. In the second experiment, vowel durations of two versions of some stimuli differed by 100 ms, to enhance the difference in syllable rhyme durations. Two subjects received the same instruction sets as in Experiment 1, and again found alignment with respect to offsets impossible. These subjects showed differences among the other instruction sets, although the direction and magnitude of the differences indicated that they had not succeeded in changing their timing criteria. The results indicate that P-center alignments are the syllable timing judgments that subjects most naturally make, and they may, indeed, be the only isochrony judgments that subjects can make reliably.

Analysis of Variance

Vowel and consonant judgments are not independent when cued by the same information.

Despite many attempts to define the major unit of speech perception, none has been generally accepted. In a unique study, Mermelstein (1978) claimed that consonants and vowels are the appropriate units because a single piece of information (duration, in this case) can be used for one distinction without affecting the other. In a replication, this apparent independence was found, instead, to reflect a lack of statistical power: The vowel and consonant judgments did interact. In another experiment, interdependence of two phonetic judgments was found in responses based on the fricative noise and the vocalic formants of a fricative-vowel syllable. These results show that each judgment made on speech signals must take into account other judgments that compete for information in the same signal. An account is proposed that takes segments as the primary units, with syllables imposing constraints on the shape they may take.

Adult

Speech perception takes precedence over nonspeech perception.

Some components of a speech signal, when made more intense, are heard simultaneously as speech and nonspeech--a form of duplex perception. At lower intensities, the speech alone is heard. Such intensity-dependent duplexity implies the existence of a phonetic mode of perception that takes precedence over auditory modes.

Adult