PubMed Health⌕ Search

Biomedical subjects

B H Repp

Publications and source records attributed to B H Repp.

At least 37 records · Page 2Linked to original sources

Detectability of duration and intensity increments in melody tones: a partial connection between music perception and performance.

Two experiments demonstrate positional variation in the relative detectability of, respectively, local temporal and dynamic perturbations in an isochronous and isodynamic sequence of melody tones, played on a computer-controlled piano. This variation may reflect listeners' expectations of expressive performance microstructure (the top-down hypothesis), or it may be due to psychoacoustic (pitch-related) stimulus factors (the bottom-up hypothesis). Percent correct scores for increments in tone duration correlated significantly with the average timing profile of pianists' expressive performances of the music, as predicted specifically by the top-down hypothesis. For intensity increments, the analogous perception-performance correlation was weak and the bottom-up factors of relative pitch height and/or direction of pitch change accounted for some of the perceptual variation. Subjects' musical training increased overall detection accuracy but did not affect the positional variation in accuracy scores in either experiment. These results are consistent with the top-down hypothesis for timing, but they favor the bottom-up hypothesis for dynamics. The perception-performance correlation for timing may also be viewed as being due to complex stimulus properties such as tonal motion and tension/relaxation that influence performers and listeners in similar ways.

Adolescent↗

Relational invariance of expressive microstructure across global tempo changes in music performance: an exploratory study.

This study addressed the question of whether the expressive microstructure of a music performance remains relationally invariant across moderate (musically acceptable) changes in tempo. Two pianists played Schumann's "Träumerei" three times at each of three tempi on a digital piano, and the performance data were recorded in MIDI format. In a perceptual test, musically trained listeners attempted to distinguish the original performances from performances that had been artificially speeded up or slowed down to the same overall duration. Accuracy in this task was barely above chance, suggesting that relational invariance was largely preserved. Subsequent analysis of the MIDI data confirmed that each pianist's characteristic timing patterns were highly similar across the three tempi, although there were statistically significant deviations from perfect relational invariance. The timing of (relatively slow) grace notes seemed relationally invariant, but selective examination of other detailed temporal features (chord asynchrony, tone overlap, pedal timing) revealed no systematic scaling with tempo. Finally, although the intensity profile seemed unaffected by tempo, a slight overall increase in intensity with tempo was observed. Effects of musical structure on expressive microstructure were large and pervasive at all levels, as were individual differences between the two pianists. For the specific composition and range of tempi considered here, these results suggest that major (cognitively controlled) temporal and dynamic features of a performance change roughly in proportion with tempo, whereas minor features tend to be governed by tempo-independent motoric constraints.

Analysis of Variance↗

Some empirical observations on sound level properties of recorded piano tones.

Preliminary to an attempt at measuring the relative intensities of overlapping tones in acoustically recorded piano music, this study investigated whether the relative peak sound levels of recorded piano tones can be reliably inferred from the levels of their two lowest harmonics, measured in the spectrum near tone onset. Acoustic recordings of single tones were obtained from two computer-controlled mechanical pianos, one upright (Yamaha MX100A Disclavier) and one concert grand (Bösendorfer 290SE), at a range of pitches and hammer velocities. Electronic recordings from a digital piano (Roland RD250S), which were free of mechanical and sound transmission factors, were included for comparison. It was found that, on all three instruments, the levels of the lowest two harmonics (in dB) near tone onset generally increased linearly with the peak root-mean-square (rms) level (in dB) as hammer velocity was varied for any given pitch. The slope of this linear function was fairly constant across midrange pitches (C2 to C6) for the first harmonic (the fundamental), but increased with pitch for the second harmonic. However, there were two sources of unpredictable variability: On the two mechanical pianos, peak rms level varied considerably across pitches, even though the strings were struck at nominally equal hammer velocities; this was probably due to the combined effects of unevenness in hammer-string interaction, soundboard response, and room acoustics. Moreover, for different pitches at equal peak rms levels, the levels of the two lowest harmonics varied substantially, even on the electronic instrument.(ABSTRACT TRUNCATED AT 250 WORDS)

Acoustics↗

Probing the cognitive representation of musical time: structural constraints on the perception of timing perturbations.

To determine whether structural factors interact with the perception of musical time, musically literate listeners were presented repeatedly with eight-bar musical excerpts, realized with physically regular timing on an electronic piano. On each trial, one or two randomly chosen time intervals were lengthened by a small amount, and the score. The resulting detection accuracy profile across all positions in each musical excerpt showed pronounced dips in places where lengthening would typically occur in an expressive (temporally modulated) performance. False alarm percentages indicated that certain tones seemed longer a priori, and these were among the ones whose actual lengthening was easiest to detect. The detection accuracy and false alarm profiles were significantly correlated with each other and with the temporal microstructure of expert performances, as measured from sound recordings by famous artists. Thus the detection task apparently tapped into listeners' musical thought and revealed their expectations about the temporal microstructure of music performance. These expectations, like the timing patterns of actual performances, derive from the cognitive representation of musical structure, as cued by a variety of systemic factors (grouping, meter, harmonic progression) and their acoustic correlates. No simple psycho-acoustic explanation of the detection accuracy profiles was evident. The results suggest that the perception of musical time is not veridical but "warped" by the structural representation. This warping may provide a natural basis for performance evaluation: expected timing patterns sound more or less regular, unexpected ones irregular. Parallels to language performance and perception are noted.

Acoustic Stimulation↗

Lexical mediation between sight and sound in speechreading.

In two experiments, we investigated whether simultaneous speech reading can influence the detection of speech in envelope-matched noise. Subjects attempted to detect the presence of a disyllabic utterance in noise while watching a speaker articulate a matching or a non-matching utterance. Speech detection was not facilitated by an audio-visual match, which suggests that listeners relied on low-level auditory cues whose perception was immune to cross-modal top-down influences. However, when the stimuli were words (Experiment 1), there was a (predicted) relative shift in bias, suggesting that the masking noise itself was perceived as more speechlike when its envelope corresponded to the visual information. This bias shift was absent, however, with non-word materials (Experiment 2). These results, which resemble earlier findings obtained with orthographic visual input, indicate that the mapping from sight to sound is lexically mediated even when, as in the case of the articulatory-phonetic correspondence, the cross-modal relationship is non-arbitrary.

Attention↗

Diversity and commonality in music performance: an analysis of timing microstructure in Schumann's "Träumerei".

This study attempts to characterize the temporal commonalities and differences among distinguished pianists' interpretations of a well-known piece, Robert Schumann's "Träumerei." Intertone onset intervals (IOIs) were measured in 28 recorded performances. These data were subjected to a variety of statistical analyses, including principal components analysis of longer stretches of music and curve fitting to series of IOIs within brief melodic gestures. Global timing patterns reflected the hierarchical grouping structure of the composition, with pronounced ritardandi at the ends of major sections and frequent expressive lengthening of accented tones within melodic gestures. Analysis of local timing patterns, particularly of within-gesture ritardandi, revealed that they often followed a parabolic timing function. The major variation in these patterns can be modeled by families of parabolas with a single degree of freedom. The grouping structure, which prescribes the location of major tempo changes, and the parabolic timing function, which represents a natural manner of executing such changes, seem to be the two major constraints under which pianists are operating. Within these constraints, there is room for much individual variation, and there are always exceptions to the rules. The striking individuality of two legendary pianists, Alfred Cortot and Vladimir Horowitz, is objectively demonstrated here, as is the relative eccentricity of several other artists.

Female↗

Perceptual restoration of a "missing" speech sound: auditory induction or illusion?

This study investigated whether the apparent completeness of the acoustic speech signal during phonemic restoration derives from a process of auditory induction (Warren, 1984) or segregation, or whether it is an auditory illusion that accompanies the completion of an abstract phonological representation. Specifically, five experiments tested the prediction of the auditory induction (segregation) hypothesis that active perceptual restoration of an [s] noise that has been replaced with an extraneous noise would use up a portion of that noise's high-frequency energy and consequently change the perceived pitch (timbre, brightness) of the extraneous noise. Listeners were required to compare the pitch of a target noise, which replaced a fricative noise in a sentence, with that of a probe noise preceding or following the speech. In the first two experiments, a significant tendency was found in favor of the auditory induction hypothesis, although the effect was small and may have been caused by variations in acoustic context. In the following three experiments, a larger variety of stimuli were used and context was controlled more carefully; this yielded negative results. Phoneme identification responses collected in the same experiments, as well as informal observations about the quality of the restored phoneme, suggested that restoration of a fricative phone distinct from the extraneous noise did not occur; rather, the spectrum of the extraneous noise itself influenced phoneme identification. These results suggest that the apparent auditory restoration which accompanies phonemic restoration is illusory, and that the schema-guided process of phoneme restoration does not interact with auditory processing.

Adult↗

Effects of preceding context on the voice-onset-time category boundary.

The category boundary on a voice-onset-time (VOT) continuum ranging from bin to pin shifts when the stimuli are preceded by different carrier phrases. By using a variety of precursor phrases, by varying the temporal interval between precursor and test word, and by selectively eliminating either voicing information or spectral structure from the precursors, the present experiments show that the context effect is caused by the presence or absence of voicing in the precursor's final segment. The effect decreases with temporal separation but persists over several seconds. In addition, the "baseline" VOT boundary for isolated test words interspersed in a test sequence shifts depending on what precursor stimuli occur in the same sequence. The perception of VOT thus seems to be sensitive to both close and distant manifestations of laryngeal activity, always in an assimilative fashion, which suggests an integrative perceptual mechanism with a long time constant.

Adult↗

Patterns of expressive timing in performances of a Beethoven minuet by nineteen famous pianists.

The timing patterns of 19 complete performances of the third movement of Beethoven's Piano Sonata op. 31, No. 3, were measured from oscillograms and analyzed statistically. One purpose of the study was to search for a timing pattern resembling the "Beethoven pulse" [Clynes, in Studies of Music Performance (Royal Academy of Music, Stockholm, 1983), pp. 76-181]. No constant pulse was found at the surface in any of the performances. Local patterns could be interpreted as evidence for an "underlying" pulse of the kind described by Clynes, but they could also derive from structural musical factors. On the whole, the artists' timing patterns served to underline the structure of the piece; lengthening at phrase boundaries and at moments of melodic/harmonic tension were the most salient features. A principal components analysis suggested that these timing variations in the Minuet could be described in terms of two orthogonal factors, one capturing mainly phrase-final lengthening, and the other reflecting phrase-internal variation as well as tempo changes. A group of musically experienced listeners evaluated the performances on a number of rating scales. Their judgments showed some significant relations to the measured timing patterns. Principal components analysis of the rating scales yielded four dimensions interpreted as force, individuality, depth, and speed. These preliminary results are encouraging for the development of more precise methods of music performance evaluation.

Humans↗

Stimulus order effects in vowel discrimination.

In same-different discrimination tasks employing isolated vowel sounds, subjects often give significantly more "different" responses to one order of two stimuli than to the other order. Cowan and Morse [J. Acoust. Soc. Am. 79, 500-507 (1986)] proposed a neutralization hypothesis to account for such effects: The first vowel in a pair is assumed to change its quality in memory in the direction of the neutral vowel, schwa. Three experiments were conducted using a variety of vowels and some initial support for the hypothesis was obtained, using a large stimulus set, but conflicting evidence with smaller stimulus sets. Rather than becoming more similar to schwa, the first vowel in a pair seems to drift toward the interior of the stimulus range employed in a given test. Several possible explanations are discussed for this tendency and its relation to presentation order effects obtained in other psychophysical paradigms is noted.

Adult↗

Acoustic properties and perception of stop consonant release transients.

This study focuses on the initial component of the stop consonant release burst, the release transient. In theory, the transient, because of its impulselike source, should contain much information about the vocal tract configuration at release, but it is usually weak in intensity and difficult to isolate from the accompanying frication in natural speech. For this investigation, a human talker produced isolated release transients of /b,d,g/ in nine vocalic contexts by whispering these syllables very quietly. He also produced the corresponding CV syllables with regular phonation for comparison. Spectral analyses showed the isolated transients to have a clearly defined formant structure, which was not seen in natural release bursts, whose spectra were dominated by the frication noise. The formant frequencies varied systematically with both consonant place of articulation and vocalic context. Perceptual experiments showed that listeners can identify both consonants and vowels from isolated transients, though not very accurately. Knowing one of the two segments in advance did not help, but when the transients were followed by a compatible synthetic, steady-state vowel, consonant identification improved somewhat. On the whole, isolated transients, despite their clear formant structure, provided only partial information for consonant identification, but no less so, it seems, than excerpted natural release bursts. The information conveyed by artificially isolated transients and by natural (frication-dominated) release bursts appears to be perceptually equivalent.

Humans↗

Effects of preceding context on discrimination of voice onset times.

When discriminating pairs of speech stimuli from an acoustic voice onset time (VOT) continuum (for example, one ranging from /ba/ to /pa/), English-speaking subjects show a characteristic performance peak in the region of the phonemic category boundary. We demonstrate that this "category boundary effect" is reduced or eliminated when the stimuli are preceded by /s/. This suppression does not seem to be due to the absence of a phonological voicing contrast for stop consonants following /s/, since it is also obtained when the /s/ terminates a preceding word and (to a lesser extent) when broadband noise is substituted for the fricative noise. The suppression is stronger, however, when the noise has the acoustic properties of a syllable-initial /s/, all else being equal. We hypothesize that these properties make the noise cohere with the following speech signal, which makes it difficult for listeners to focus on the VOT differences to be discriminated.

Adult↗

Perception of the [m]-[n] distinction in VC syllables.

This study complements earlier experiments on the perception of the [m]-[n] distinction in CV syllables [B. H. Repp, J. Acoust. Soc. Am. 79, 1987-1999 (1986); B. H. Repp, J. Acoust. Soc. Am. 82, 1525-1538 (1987)]. Six talkers produced VC syllables consisting of [m] or [n] preceded by [i, a, u]. In listening experiments, these syllables were truncated from the beginning and/or from the end, or waveform portions surrounding the point of closure were replaced with noise, so as to map out the distribution of the place of articulation information for consonant perception. These manipulations revealed that the vocalic formant transitions alone conveyed about as much place of articulation information as did the nasal murmur alone, and both signal portions were about as informative in VC as in CV syllables. Nevertheless, full VC syllables were less accurately identified than full CV syllables, especially in female speech. The reason for this was hypothesized to be the relative absence of a salient spectral change between the vowel and the murmur in VC syllables. This hypothesis was supported by the relative ineffectiveness of two additional manipulations meant to disrupt the perception of relational spectral information (channel separation or temporal separation of vowel and murmur) and by subjects' poor identification scores for brief excerpts including the point of maximal spectral change. While, in CV syllables, the abrupt spectral change from the murmur to the vowel provides important additional place of articulation information, for VC syllables it seems as if the format transitions in the vowel and the murmur spectrum functioned as independent cues.

Adult↗

Detectability of words and nonwords in two kinds of noise.

Recent models of speech perception emphasize the possibility of interactions among different processing levels. There is evidence that the lexical status of an utterance (i.e., whether it is a meaningful word or not) may influence earlier stages of perceptual analysis. To test how far down such "top-down" influences might penetrate, an investigation was conducted to determine whether there is a difference in detectability of words and nonwords masked by amplitude-modulated or unmodulated broadband noise. The results were negative, suggesting either that the stages of perceptual analysis engaged in the detection task are impermeable to lexical top-down effects, or that the lexical level was not sufficiently activated to have any facilitative effect on perception.

Humans↗

The sound of two hands clapping: an exploratory study.

Clapping is a little-studied human activity that may be viewed either as a form of communicative group behavior (applause) or as an individual sound-generating activity involving two "articulators"--the hands. The latter aspect was explored in this pilot study by means of acoustical analyses and perceptual experiments. Principal components analysis of 20 subjects' average clap spectra yielded several dimensions of interindividual variation that were related to observed hand configuration. This relationship emerged even more clearly in a similar analysis of a single clapper's deliberately varied productions. In perception experiments, subjects proved sensitive to spectral properties of claps: For a single clapper, at least, listeners were able to judge hand configuration with good accuracy. Besides providing some general information on individual variations in clapping, the present results support the general hypothesis that sound emanating from a natural source informs listeners about the changing states of the source mechanism.

Adult↗

On the possible role of auditory short-term adaptation in perception of the prevocalic [m]-[n] contrast.

Acoustic information about the place of articulation of a prevocalic nasal consonant is distributed over two distinct signal portions, the nasal murmur and the onset of the following vowel. The spectral properties of these signal portions are perceptually important, as is their relationship (the pattern of spectral change). A series of experiments was conducted to investigate to what extent relational place of articulation information derives from a peripheral auditory interaction, viz., short-term adaptation caused by the murmur. Experimental manipulations intended to disrupt the effects of such adaptation included separation of the murmur and the vowel by intervals of silence, presentation to different ears, and reversal of order. Other tests of the possible role of adaptation included manipulation of murmur duration, murmur-vowel cross splicing, and high-pass filtering of the excised vowel onset. While the results of several experiments were compatible with the peripheral adaptation hypothesis, others did not support it. An alternative hypothesis, that the manner cues provided by the murmur are crucial for accurate place judgments, was also discredited. It was concluded that, at least under good listening conditions, the perception of spectral relationships does not depend on peripheral auditory enhancement and probably rests on a central comparison process.

Acclimatization↗

Perception of the [m]-[n] distinction in CV syllables.

The contribution of the nasal murmur and the vocalic formant transitions to perception of the [m]-[n] distinction in utterance-initial position preceding [i,a,u] was investigated, extending the recent work of Kurowski and Blumstein [J. Acoust. Soc. Am. 76, 383-390 (1984)]. A variety of waveform-editing procedures were applied to syllables produced by six different talkers. Listeners' judgments of the edited stimuli confirmed that the nasal murmur makes a significant contribution to place of articulation perception. Murmur and transition information appeared to be integrated at a genuinely perceptual, not an abstract cognitive, level. This was particularly evident in [-i] context, where only the simultaneous presence of murmur and transition components permitted accurate place of articulation identification. The perceptual information seemed to be purely relational in this case. It also seemed to be context specific, since the spectral change from the murmur to the vowel onset did not follow an invariant pattern across front and back vowels.

Female↗