PubMed Health⌕ Search

Biomedical subjects

W Strange

Publications and source records attributed to W Strange.

At least 19 recordsLinked to original sources

Effects of consonantal context on perceptual assimilation of American English vowels by Japanese listeners.

This study investigated the extent to which adult Japanese listeners' perceived phonetic similarity of American English (AE) and Japanese (J) vowels varied with consonantal context. Four AE speakers produced multiple instances of the 11 AE vowels in six syllabic contexts /b-b, b-p, d-d, d-t, g-g, g-k/ embedded in a short carrier sentence. Twenty-four native speakers of Japanese were asked to categorize each vowel utterance as most similar to one of 18 Japanese categories [five one-mora vowels, five two-mora vowels, plus/ei, ou/ and one-mora and two-mora vowels in palatalized consonant CV syllables, C(j)a(a), C(j)u(u), C(j)o(o)]. They then rated the "category goodness" of the AE vowel to the selected Japanese category on a seven-point scale. None of the 11 AE vowels was assimilated unanimously to a single J response category in all context/speaker conditions; consistency in selecting a single response category ranged from 77% for /eI/ to only 32% for /ae/. Median ratings of category goodness for modal response categories were somewhat restricted overall, ranging from 5 to 3. Results indicated that temporal assimilation patterns (judged similarity to one-mora versus two-mora Japanese categories) differed as a function of the voicing of the final consonant, especially for the AE vowels, /see text/. Patterns of spectral assimilation (judged similarity to the five J vowel qualities) of /see text/ also varied systematically with consonantal context and speakers. On the basis of these results, it was predicted that relative difficulty in the identification and discrimination of AE vowels by Japanese speakers would vary significantly as a function of the contexts in which they were produced and presented.

Adolescent↗

Context-independent dynamic information for the perception of coarticulated vowels.

Most investigators agree that the acoustic information for American English vowels includes dynamic (time-varying) parameters as well as static "target" information contained in a single cross section of the syllable. Using the silent-center (SC) paradigm, the present experiment examined the case in which the initial and final portions of stop consonant-vowel-stop consonant (CVC) syllables containing the same vowel but different consonants were recombined into mixed-consonant SC syllables and presented to listeners for vowel identification. Ten vowels were spoken in six different syllables, /b Vb, bVd, bVt, dVb, dVd, dVt/, embedded in a carrier sentence. Initial and final transitional portions of these syllables were cross-matched in: (1) silent-center syllables with original syllable durations (silences) preserved (mixed-consonant SC condition) and (2) mixed-consonant SC syllables with syllable duration equated across the ten vowels (fixed duration mixed-consonant SC condition). Vowel-identification accuracy in these two mixed consonant SC conditions was compared with performance on the original SC and fixed duration SC stimuli, and in initial and final control conditions in which initial and final transitional portions were each presented alone. Vowels were identified highly accurately in both mixed-consonant SC and original syllable SC conditions (only 7%-8% overall errors). Neutralizing duration information led to small, but significant, increases in identification errors in both mixed-consonant and original fixed-duration SC conditions (14%-15% errors), but performance was still much more accurate than for initial and finals control conditions (35% and 52% errors, respectively). Acoustical analysis confirmed that direction and extent of formant change from initial to final portions of mixed-consonant stimuli differed from that of original syllables, arguing against a target + offglide explanation of the perceptual results. Results do support the hypothesis that temporal trajectories specifying "style of movement" provide information for the differentiation of American English tense and lax vowels, and that this information is invariant over the place of articulation and voicing of the surrounding stop consonants.

Adult↗

Perception of dynamic information for vowels in syllable onsets and offsets.

It has been demonstrated using the "silent-center" (SC) syllable paradigm that there is sufficient information in syllable onsets and offsets, taken together, to support accurate identification of vowels spoken in both citation-form syllables and syllables spoken in sentence context. Using edited natural speech stimuli, the present study examined the identification of American English vowels when increasing amounts of syllable onsets alone or syllable offsets alone were presented in their original sentence context. The stimuli were /d/-vowel-/d/ syllables spoken in a short carrier sentence by a male speaker. Listeners attempted to identify the vowels in experimental conditions that differed in the number of pitch periods presented and whether the pitch periods were from syllable onsets or syllable offsets. In general, syllable onsets were more informative than syllable offsets, although neither onsets nor offsets alone specified vowel identity as well as onsets and offsets together (SC syllables). Vowels differed widely in ease of identification; the diphthongized long vowels /e/, /ae/, /o/ were especially difficult to identify from syllable offsets. Identification of vowels as "front" or "back" was accurate, even from short samples of the syllable; however, vowel "height" was quite difficult to determine, again, especially from syllable offsets. The results emphasize the perceptual importance of time-varying acoustic parameters, which are the direct consequence of the articulatory dynamics involved in producing syllables.

Adult↗

Dynamic specification of coarticulated German vowels: perceptual and acoustical studies.

To examine the generality of Strange's Dynamic Specification Theory of vowel perception, two perceptual experiments investigated whether dynamic (time-varying) acoustic information about vowel gestures was critical for identification of coarticulated vowels in German, a language without diphthongization. The perception by native North German (NG) speakers of electronically modified /dVt/ syllables produced in carrier sentences was assessed using the "silent-center" paradigm. The relative efficacy of static target information, dynamic spectral information (defined over syllable onsets and offsets together), and intrinsic vowel length was investigated in listening conditions in which the centers (silent-center conditions) or the onsets and offsets (vowel-center conditions) of the syllables were silenced. Listeners correctly identified most vowels in silent-center syllables and in vowel-center stimuli when both conditions included information about intrinsic vowel length. When duration information was removed, errors increased significantly, but performance was relatively better for silent-center syllables than for vowel-center stimuli. Acoustical analyses of the effects of coarticulation on target formant frequencies, vocalic duration, and dynamic spectro-temporal patterns in the stimulus materials were performed to elucidate the nature of the dynamic spectral information. In comparison with vowels produced in citation from /hVt/ syllables by the same speaker, the coarticulated /dVt/ utterances showed considerable "target undershoot" of formant frequencies and reduced duration differences between tense and lax vowel pairs. This suggests that both static spectral cues and relative duration information for NG vowels may not remain perceptually distinctive in continuous speech. Analysis of formant movement within syllable nuclei corroborated descriptions of German vowels as monophthongal. However, an analysis of first formant temporal trajectories revealed distinct patterns for tense and lax vowels that could be used by listeners to disambiguate coarticulated NG vowels.

Adult↗

Vowel identification in mixed-speaker silent-center syllables.

Strange [J. Acoust. Soc. Am. 85, 2135-2153 (1989b)] has demonstrated that there is sufficient information in the onsets and offsets of syllables, spoken in sentence context, to provide accurate identification of the vowel in the syllable. Verbrugge and Rakerd [Language Speech 29, 39-57 (1986)] and Andruski and Nearey [J. Acoust. Soc. Am. 91, 390-490 (1992)] have shown that such information is present in citation-form syllables even when the syllables begin with one speaker and end with another. These studies of "hybrid syllables," however, reported relatively high error rates. In a perceptual experiment using /dVd/ syllables spoken in sentence context by a male and a female speaker, relatively low error rates were obtained for both "silent-center" syllables and "hybrid silent-center" syllables. It was concluded that the information specified over syllable onsets and offsets together for identification of vowels is speaker independent and that it was sufficient in most cases to specify the vowel. The acoustic patterns, represented as functions of log (F2/F1) over time, revealed potentially useful dynamic acoustic characteristics of co-articulated vowels.

Adult↗

Evolving theories of vowel perception.

Research on the perception of vowels in the last several years has given rise to new conceptions of vowels as articulatory, acoustic, and perceptual events. Starting from a "simple" target model in which vowels were characterized articulatorily as static vocal tract shapes and acoustically as points in a first and second formant (F1/F2) vowel space, this paper briefly traces the evolution of vowel theory in the 1970s and 1980s in two directions. (1) Elaborated target models represent vowels as target zones in perceptual spaces whose dimensions are specified as formant ratios. These models have been developed primarily to account for perceivers' solution of the "speaker normalization" problem. (2) Dynamic specification models emphasize the importance of formant trajectory patterns in specifying vowel identity. These models deal primarily with the problem of "target undershoot" associated with the coarticulation of vowels with consonants in natural speech and with the issue of "vowel-inherent spectral change" or diphthongization of English vowels. Perceptual studies are summarized that motivate these theoretical developments.

Auditory Threshold↗

Dynamic specification of coarticulated vowels spoken in sentence context.

According to a dynamic specification account, coarticulated vowels are identified on the basis of time-varying acoustic information, rather than solely on the basis of "target" information contained within a single spectral cross section of an acoustic syllable. Three experiments utilizing digitally segmented portions of consonant-vowel-consonant (CVC) syllables spoken rapidly in a carrier sentence were designed to examine the relative contribution of (1) target information available in vocalic nuclei, (2) intrinsic duration information specified by syllable length, and (3) dynamic spectral information defined over syllable onsets and offsets. In experiments 1 and 2, vowels produced in three consonantal contexts by an adult male were examined. Results showed that vowels in silent-center (SC) syllables (in which vocalic nuclei were attentuated to silence leaving initial and final transitional portions in their original temporal relationship) were perceived relatively accurately, although not as well as unmodified syllables (experiment 1); random versus blocked presentation of consonantal contexts did not affect performance. Error rates were slightly greater for vowels in SC syllables in which intrinsic duration differences were neutralized by equating the duration of silent intervals between initial and final transitional portions. However, performance was significantly better than when only initial transitions or final transitions were presented alone (experiment 2). Experiment 3 employed CVC stimuli produced by another adult male, and included six consonantal contexts. Both SC syllables and excised syllable nuclei with appropriate intrinsic durations were identified no less accurately than unmodified controls. Neutralizing duration differences in SC syllables increased identification errors only slightly, while truncating excised syllable nuclei yielded a greater increase in errors. These results demonstrate that time-varying information is necessary for accurate identification of coarticulated vowels. Two hypotheses about the nature of the dynamic information specified over syllable onsets and offsets are discussed.

Adolescent↗

Trading relations in the perception of /r/-/l/ by Japanese learners of English.

The role of language-specific factors in phonetically based trading relations was examined by assessing the ability of 20 native Japanese speakers to identify and discriminate stimuli of two synthetic /r/-/l/ series that varied temporal and spectral parameters independently. Results of forced-choice identification and oddity discrimination tasks showed that the nine Japanese subjects who were able to identify /r/ and /l/ reliably demonstrated a trading relation similar to that of Americans. Discrimination results reflected the perceptual equivalence of temporal and spectral parameters. Discrimination by the 11 Japanese subjects who were unable to identify the /r/-/l/ series differed significantly from the skilled Japanese subjects and native English speakers. However, their performance could not be predicted on the basis of acoustic dissimilarity alone. These results provide evidence that the trading relation between temporal and spectral cues for the /r/-/l/ contrast is not solely attributable to general auditory or language-universal phonetic processing constraints, but rather is also a function of phonemic processes that can be modified in the course of learning a second language.

Adolescent↗

Perceptual equivalence of acoustic cues that differentiate /r/ and /l/.

The perceptual effects of orthogonal variations in two acoustic parameters which differentiate American English prevocalic /r/ and /l/ were examined. A spectral cue (frequency onset and transition of F2 and F3) and a temporal cue (relative duration of initial steady state and transition of F1) were varied in synthetic versions of "rock" and "lock." Four temporal variations in each of ten stimuli of a spectral-cue continuum were generated. Phonetic identification and oddity discrimination tasks with the four series showed systematic displacement of perceptual boundaries and discrimination peaks, thus reflecting a trading relation between the two cues. The perceptual equivalence of spectral and temporal cues was investigated by comparing the accuracy of discrimination of three types of stimulus comparisons: phonetically facilitating two-cue pairs, one-cue pairs, and phonetically conflicting two-cue pairs. As predicted, discrimination accuracy was ordered: Facilitating cues greater than one-cue greater than conflicting cues, indicating that perceivers discriminated on the basis of an integrated phonetic percept.

Acoustic Stimulation↗

Perception and production of approximant consonants by normal and articulation-delayed preschool children.

Disagreement exists concerning the relationship between the perception of phonetic contrasts and their production by both normal and articulation-delayed children. The perception of three approximant consonant contrasts (/w/-/r/, /w/-/l/, /r/-/l/) was examined in two groups of 3-year-old children: normal children who did and did not articulate /r/ and /l/ correctly and articulation-delayed children who misarticulated /r/ and /l/. Perception was assessed in a two-choice forced-choice identification task in which the subjects heard a word and pushed a button lighting a picture corresponding to the word. In general, normally developing children were highly accurate in their perception of all three contrasts, but there was more variability in /w/-/l/ perceptual performance among the children who neutralized the /w/-/l/ contrast. Articulation-delayed children displayed a wider range of production patterns and were more variable in their perceptual performance than normally developing children. Results suggest than normally developing children learn to perceive approximant contrasts prior to 3 years of age. However, some but not all articulation-delayed 3-year-old children may still make errors in the perception of approximants.

Articulation Disorders↗

Dynamic specification of coarticulated vowels.

An adequate theory of vowel perception must account for perceptual constancy over variations in the acoustic structure of coarticulated vowels contributed by speakers, speaking rate, and consonantal context. We modified recorded consonant-vowel-consonant syllables electronically to investigate the perceptual efficacy of three types of acoustic information for vowel identification: (1) static spectral "targets," (2) duration of syllabic nuclei, and (3) formant transitions into and out of the vowel nucleus. Vowels in /b/-vowel-/b/ syllables spoken by one adult male (experiment 1) and by two females and two males (experiment 2) served as the corpus, and seven modified syllable conditions were generated in which different parts of the digitized waveforms of the syllables were deleted and the temporal relationships of the remaining parts were manipulated. Results of identification tests by untrained listeners indicated that dynamic spectral information, contained in initial and final transitions taken together, was sufficient for accurate identification of vowels even when vowel nuclei were attenuated to silence. Furthermore, the dynamic spectral information appeared to be efficacious even when durational parameters specifying intrinsic vowel length were eliminated.

Female↗

Cross-language study of perception of the oral-nasal distinction.

To investigate the effect of linguistic experience on the perception of the oral-nasal distinction in vowels, Hindi and American English speakers were tested on identification and discrimination of four speechlike series generated by articulatory synthesis. In experiment I, no language-group differences were found in the discrimination of either a consonant series, [ba-ma] (a phonemic contrast in both languages), or a vowel series, [ba-bã] (phonemic only for the Hindi speakers). The vowel results were due to floor effects which obscured differences across language groups. Experiment II examined perception of two modified vowel series, one which increased the interval between members of discrimination pairs and one that extended the range of velar port opening. Cross-language differences in discrimination were found. Hindi perception of the oral-nasal distinction was categorical. English speakers' perception of the vowel series was more continuous. They accurately discriminated differences not only across categories, but also within the oral category. These findings indicate that linguistic experience can influence listeners' perception of vowels, but the effect is different from that shown for consonants.

Cross-Cultural Comparison↗

Perceptual constancy of vowels in rapid speech.

In three experiments, we investigated the role of extrasyllabic speech context in the identification of /t/-vowel-/t/ syllables spoken at normal and rapid rates of articulation. Syllables spoken at the normal rate were identified with at least 95% accuracy regardless of the context in which they were presented. However, in experiment I, intrinsically long vowels in rapidly articulated syllables were identified with greater accuracy when the syllables were presented in rapid-rate sentence context than when they were presented in isolation or in normal-rate sentence context. Experiment II revealed that part of the performance deficit for excised isolated rapid syllables was attributable to sequential order effects among test items, and to listeners' assumptions about whether the test syllables were or were not produced in isolation. In experiment III, test syllables were presented in the context of extracted portions of the rapid carrier sentences to determine the locus of the extrasyllabic information for vowel identity. The presence of the stressed syllable which immediately followed the test syllables in the carrier sentence was especially important for accurate identification of intrinsically long vowels in rapidly articulated syllables. Results are discussed in terms of intrinsic timing theories of speech production.

Cues↗

Task variables in the study of vowel perception.

Vowels produced in isolation by six speakers were identified less accurately than vowels coarticulated in /k/-vowel-/k/ syllables, whether tested by a key word task or by a rhyming task. Performance by naive listeners in the two tasks was highly correlated.

Humans↗

Identification of coarticulated vowels.

Previous explanations of vowel perception held that the most definitive information for vowel identity is the relatively constant formant frequencies in the steady-state portions of vowels. Perceptual studies indicate, however, that vowels spoken in syllables with labial stop consonants are identified more accurately than vowels spoken in isolation. The present study investigated the nature and scope of this consonantal context advantage in the perception of ten American English vowels spoken by adult male and female speakers. Vowels in /p/-vowel-/p/, /b/-vowel-/b/, /k/-vowel-/k/, /k/, /k/-vowel, and vowel-/k/ syllables were identified much more accurately than isolated vowels. This is consistent with the hypothesis that dynamic acoustic information due to the coarticulation in syllables is important for vowel identification. Identification of vowels in /g/-vowel-/g/, /g/-vowel, and vowel-/g/ syllables was not better than isolated vowels and was significantly poorer than for other consonantal contexts. Acoustical analyses were performed to determine whether poor production of vowels could account for perceptual errors. Misproduced vowel targets could not account for the overall pattern of identification performance. Phonological factors were also considered but were found to be inadequate to account fully for the results.

Adult↗