PubMed Health⌕ Search

Biomedical subjects

Paavo Alku

Publications and source records attributed to Paavo Alku.

At least 19 recordsLinked to original sources

Changes in objective acoustic measurements and subjective voice complaints in call center customer-service advisors during one working day.

SUMMARY: The aim of this study was to investigate how different acoustic parameters, extracted both from speech pressure waveforms and glottal flows, can be used in measuring vocal loading in modern working environments and how these parameters reflect the possible changes in the vocal function during a working day. In addition, correlations between objective acoustic parameters and subjective voice symptoms were addressed. The subjects were 24 female and 8 male customer-service advisors, who mainly use telephone during their working hours. Speech samples were recorded from continuous speech four times during a working day and voice symptom questionnaires were completed simultaneously. Among the various objective parameters, only F0 resulted in a statistically significant increase for both genders. No correlations between the changes in objective and subjective parameters appeared. However, the results encourage researchers within the field of occupational voice use to apply versatile measurement techniques in studying occupational voice loading.

Adult↗

Group intervention changes brain activity in bilingual language-impaired children.

This investigation assessed the effectiveness of a phonological intervention program on the brain functioning of bilingual Finnish 6- to 7-year-old preschool children diagnosed with specific language impairment (SLI). The intervention program was implemented by preschool teachers to small groups of children including children with SLI. A matched group of other bilingual children with SLI received a physical exercise program and served as a control group. Auditory evoked magnetic fields were measured before and after the intervention with an oddball paradigm. The brain activity recordings were followed by a behavioral discrimination test. Our results show that, in children with SLI, the positive intervention effect is reflected in plastic changes in the brain activity of the left and right auditory cortices.

Adaptation, Physiological↗

Emotions in vowel segments of continuous speech: analysis of the glottal flow using the normalised amplitude quotient.

Emotions in short vowel segments of continuous speech were analysed using inverse filtering and a recently developed glottal flow parameter, the normalised amplitude quotient (NAQ). Simulated emotion portrayals were produced by 9 professional stage actors. Separated /a:/ vowel segments were inverse filtered and parameterized using NAQ. Statistical analyses showed significant differences among most of the emotions studied. Results also demonstrated clear gender differences. Inverse filtering, together with NAQ, was shown to be a promising method for the analysis of emotional content in continuous speech.

Adult↗

Comparison of two inverse filtering methods in parameterization of the glottal closing phase characteristics in different phonation types.

SUMMARY: Inverse filtering (IF) is a common method used to estimate the source of voiced speech, the glottal flow. This investigation aims to compare two IF methods: one manual and the other semiautomatic. Glottal flows were estimated from speech pressure waveforms of six female and seven male subjects producing sustained vole /a/ in breathy, normal, and pressed phonation. The closing phase characteristics of the glottal pulse were parameterized using two time-based parameters: the closing quotient (C1Q) and the normalized amplitude quotient (NAQ). The information given by these two parameters indicates a strong correlation between the two IF methods. The results are encouraging in showing that the parameterization of the voice source in different speech sounds can be performed independently of the technique used for inverse filtering.

Adult↗

Mismatch negativity (MMN) elicited by changes in phoneme length: a cross-linguistic study.

Speech sounds representing different phonetic categories are typically easier to discriminate than sounds belonging to the same category. This phenomenon is referred to as the phoneme boundary effect. We aimed to determine whether, at neural level, this effect is indeed due to crossing the phoneme boundary. The mismatch negativity (MMN) brain response was measured for across- and within-category changes in Finnish phoneme length in native speakers and second-language users of Finnish as well as non-Finnish-speaking subjects. The results showed that the MMN amplitude was enhanced in the native speakers in comparison with the two non-native groups which, in turn, did not differ from each other in MMN amplitude. The response pattern to across- and within-category changes, however, was the same in all groups regardless of whether or not they had the phoneme categories. Thus, the responses could not be determined by crossing the phoneme boundary. Rather, the enhancement of MMN amplitude in the native speakers is likely to be due to the activation of native-language phonetic prototypes. The second-language users, however, did not seem to have automatic access to Finnish prototypes.

Analysis of Variance↗

Emotions in [a]: a perceptual and acoustic study.

The aim of this investigation is to study how well voice quality conveys emotional content that can be discriminated by human listeners and the computer. The speech data were produced by nine professional actors (four women, five men). The speakers simulated the following basic emotions in a unit consisting of a vowel extracted from running Finnish speech: neutral, sadness, joy, anger, and tenderness. The automatic discrimination was clearly more successful than human emotion recognition. Human listeners thus apparently need longer speech samples than vowel-length units for reliable emotion discrimination than the machine, which utilizes quantitative parameters effectively for short speech samples.

Adult↗

The role of F3 in the vocal expression of emotions.

The present study investigates the role of F3 in the perception of valence of emotional expressions by using a vowel [a:] with different F3 values: the original, one with F3 either lowered or raised by 30% in frequency, and one with F3 removed. The vowel [a:] was extracted from the simulated emotions, inverse filtered and manipulated. The resulting 12 synthesized samples were randomized and presented to 30 listeners who evaluated the valence (positiveness/negativeness) of the expressions. The vowel with raised F3 was perceived more often as positive than the sample with original (p = 0.063), lowered (p = 0.006) or removed F3 (p = 0.066). F3 may affect perception of valence if the signal has sufficient energy in high frequency range.

Adult↗

Subglottal pressure and normalized amplitude quotient variation in classically trained baritone singers.

The subglottal pressure (Ps) and voice source characteristics of five professional baritone singers have been analyzed and the normalized amplitude quotient (NAQ), defined as the ratio between peak-to-peak pulse amplitude and the negative peak of the differentiated flow glottogram and normalized with respect to the period time, was used as an estimate of glottal adduction. The relationship between Ps and NAQ has been investigated in female subjects in two earlier studies. One of these revealed NAQ differences between both singing styles and phonation modes, and the other, based on register differences in female musical theatre singers, showed that NAQ differed between registers for the same Ps value. These studies thus suggest that NAQ and its variation with Ps represent a useful parameter in the analysis of voice source characteristics. The present study aims at increasing our knowledge of the NAQ parameter further by finding out how it varies with pitch and Ps in professional classically trained baritone singers, singing at high and low pitch (278 Hz and 139 Hz, respectively). Ten equally spaced Ps values were selected from three takes of the syllable [pae:], initiated at maximum vocal loudness and repeated with a continuously decreasing vocal loudness. The vowel sounds following the selected Ps peaks were inverse filtered. Data on peak-to-peak pulse amplitude, maximum flow declination rate and NAQ are presented.

Adult↗

Occupational voice complaints and objective acoustic measurements-do they correlate?

To enable the development of appropriate diagnostics and treatment for occupational voice disorders, this study addresses connections between subjective voice complaints and objective observations. The subjects of this study were 24 female customer advisors, who mainly use the telephone during their working hours. During one working day, at four different times, speech samples covering 20 minutes of telephone conversation by the customer service advisors (CSAs) were recorded. In addition, the CSAs filled in a questionnaire (visual analogue scale) concerning their voice problems. To represent the vocal symptoms three variables were used: vocal fatigue, hoarseness and a general sum-variable. A 5-minute sample was taken from recordings for further analyses. This included fundamental frequency, sound pressure level, alpha ratio (the ratio between the spectral energy below and above 1000 Hz) and number of vocal fold vibrations. In the objective acoustic measurements, it was found that fundamental frequency (F0) rose significantly during the working day. Also the self-reported voice symptoms increased significantly during the working day. However, correlations between vocal symptoms and acoustic measures were not found.

Adult↗

An amplitude quotient based method to analyze changes in the shape of the glottal pulse in the regulation of vocal intensity.

This study presents an approach to visualizing intensity regulation in speech. The method expresses a voice sample in a two-dimensional space using amplitude-domain values extracted from the glottal flow estimated by inverse filtering. The two-dimensional presentation is obtained by expressing a time-domain measure of the glottal pulse, the amplitude quotient (AQ), as a function of the negative peak amplitude of the flow derivative (d(peak)). The regulation of vocal intensity was analyzed with the proposed method from voices varying from extremely soft to very loud with a SPL range of approximately 55 dB. When vocal intensity was increased, the speech samples first showed a rapidly decreasing trend as expressed on the proposed AQ-d(peak) graph. When intensity was further raised, the location of the samples converged toward a horizontal line, the asymptote of a hypothetical hyperbola. This behavior of the AQ-d(peak) graph indicates that the intensity regulation strategy changes from laryngeal to respiratory mechanisms and the method chosen makes it possible to quantify how control mechanisms underlying the regulation of vocal intensity change gradually between the two means. The proposed presentation constitutes an easy-to-implement method to visualize the function of voice production in intensity regulation because the only information needed is the glottal flow wave form estimated by inverse filtering the acoustic speech pressure signal.

Adult↗

Estimation of the voice source from speech pressure signals: evaluation of an inverse filtering technique using physical modelling of voice production.

OBJECTIVE: The goal of the study is to use physical modelling of voice production to assess the performance of an inverse filtering method in estimating the glottal flow from acoustic speech pressure signals. METHODS: An automatic inverse filtering method is presented, and speech pressure signals are generated using physical modelling of voice production so as to obtain test vowels with a known shape of the glottal excitation waveform. The speech sounds produced consist of 4 different vowels, each with 10 different values of the fundamental frequency. Both the original glottal flows given by physical modelling and their estimates computed by inverse filtering were parametrised with two robust voice source parameters: the normalized amplitude quotient and the difference (in decibels) between the levels of the first and second harmonics. RESULTS: The results show that for both extracted parameters the error introduced by inverse filtering was, in general, small. The effect of the distortion caused by inverse filtering on the parameter values was clearly smaller than the change in the corresponding parameters when the phonation type was altered. The distortion was largest for high-pitched vowels with the lowest value of the first formant. CONCLUSIONS: The study shows that the proposed inverse filtering technique combined with the extracted parameters constitutes a voice source analysis tool that is able to measure the voice source dynamics automatically with satisfactory accuracy.

Glottis↗

Neuromagnetic recordings reveal the temporal dynamics of auditory spatial processing in the human cortex.

In an attempt to delineate the assumed 'what' and 'where' processing streams, we studied the processing of spatial sound in the human cortex by using magnetoencephalography in the passive and active recording conditions and two kinds of spatial stimuli: individually constructed, highly realistic spatial (3D) stimuli and stimuli containing interaural time difference (ITD) cues only. The auditory P1m, N1m, and P2m responses of the event-related field were found to be sensitive to the direction of sound source in the azimuthal plane. In general, the right-hemispheric responses to spatial sounds were more prominent than the left-hemispheric ones. The right-hemispheric P1m and N1m responses peaked earlier for sound sources in the contralateral than for sources in the ipsilateral hemifield and the peak amplitudes of all responses reached their maxima for contralateral sound sources. The amplitude of the right-hemispheric P2m response reflected the degree of spatiality of sound, being twice as large for the 3D than ITD stimuli. The results indicate that the right hemisphere is specialized in the processing of spatial cues in the passive recording condition. Minimum current estimate (MCE) localization revealed that temporal areas were activated both in the active and passive condition. This initial activation, taking place at around 100 ms, was followed by parietal and frontal activity at 180 and 200 ms, respectively. The latter activations, however, were specific to attentional engagement and motor responding. This suggests that parietal activation reflects active responding to a spatial sound rather than auditory spatial processing as such.

Acoustic Stimulation↗

The discrimination of and orienting to speech and non-speech sounds in children with autism.

The present study aimed to find out how different stages of cortical auditory processing (sound encoding, discrimination, and orienting) are affected in children with autism. To this end, auditory event-related potentials (ERP) were studied in 15 children with autism and their controls. Their responses were recorded for pitch, duration, and vowel changes in speech stimuli, and for corresponding changes in the non-speech counterparts of the stimuli, while the children watched silent videos and ignored the stimuli. The responses to sound repetition were diminished in amplitude in the children with autism, reflecting impaired sound encoding. The mismatch negativity (MMN), an ERP indexing sound discrimination, was enhanced in the children with autism as far as pitch changes were concerned. This is consistent with earlier studies reporting auditory hypersensitivity and good pitch-processing abilities, as well as with theories proposing enhanced perception of local stimulus features in individuals with autism. The discrimination of duration changes was impaired in these children, however. Finally, involuntary orienting to sound changes, as reflected by the P3a ERP, was more impaired for speech than non-speech sounds in the children with autism, suggesting deficits particularly in social orienting. This has been proposed to be one of the earliest symptoms to emerge, with pervasive effects on later development.

Acoustic Stimulation↗

Disentangling the effects of phonation and articulation: hemispheric asymmetries in the auditory N1m response of the human brain.

BACKGROUND: The cortical activity underlying the perception of vowel identity has typically been addressed by manipulating the first and second formant frequency (F1 & F2) of the speech stimuli. These two values, originating from articulation, are already sufficient for the phonetic characterization of vowel category. In the present study, we investigated how the spectral cues caused by articulation are reflected in cortical speech processing when combined with phonation, the other major part of speech production manifested as the fundamental frequency (F0) and its harmonic integer multiples. To study the combined effects of articulation and phonation we presented vowels with either high (/a/) or low (/u/) formant frequencies which were driven by three different types of excitation: a natural periodic pulseform reflecting the vibration of the vocal folds, an aperiodic noise excitation, or a tonal waveform. The auditory N1m response was recorded with whole-head magnetoencephalography (MEG) from ten human subjects in order to resolve whether brain events reflecting articulation and phonation are specific to the left or right hemisphere of the human brain. RESULTS: The N1m responses for the six stimulus types displayed a considerable dynamic range of 115-135 ms, and were elicited faster (approximately 10 ms) by the high-formant /a/ than by the low-formant /u/, indicating an effect of articulation. While excitation type had no effect on the latency of the right-hemispheric N1m, the left-hemispheric N1m elicited by the tonally excited /a/ was some 10 ms earlier than that elicited by the periodic and the aperiodic excitation. The amplitude of the N1m in both hemispheres was systematically stronger to stimulation with natural periodic excitation. Also, stimulus type had a marked (up to 7 mm) effect on the source location of the N1m, with periodic excitation resulting in more anterior sources than aperiodic and tonal excitation. CONCLUSION: The auditory brain areas of the two hemispheres exhibit differential tuning to natural speech signals, observable already in the passive recording condition. The variations in the latency and strength of the auditory N1m response can be traced back to the spectral structure of the stimuli. More specifically, the combined effects of the harmonic comb structure originating from the natural voice excitation caused by the fluctuating vocal folds and the location of the formant frequencies originating from the vocal tract leads to asymmetric behaviour of the left and right hemisphere.

Acoustic Stimulation↗

Left-hemispheric brain activity reflects formant transitions in speech sounds.

Connected speech is characterized by formant transitions whereby formant frequencies change over time. Here, using magneto-encephalography, we investigated the cortical activity in 10 participants in response to constant-formant vowels and diphthongs with formant transitions. All the stimuli elicited prominent auditory N100m responses, but the formant transitions resulted in latency modulations specific to the left hemisphere. Following the elicitation of the N100m, cortical activity shifted some 10 mm towards anterior brain areas. This late activity resembled the N400m, typically obtained with more complex utterances such as words and/or sentences. Thus, the present study demonstrates how magnetoencephalography can be used to investigate the spatiotemporal evolution in cortical activity related to the various stages of the processing of speech.

Acoustic Stimulation↗

Spatial processing in human auditory cortex: the effects of 3D, ITD, and ILD stimulation techniques.

Here, the perception of auditory spatial information as indexed by behavioral measures is linked to brain dynamics as reflected by the N1m response recorded with whole-head magnetoencephalography (MEG). Broadband noise stimuli with realistic spatial cues corresponding to eight direction angles in the horizontal plane were constructed via custom-made, individualized binaural recordings (BAR) and generic head-related transfer functions (HRTF). For comparison purposes, stimuli with impoverished acoustical cues were created via interaural time and level differences (ITDs and ILDs) and their combinations. MEG recordings in ten subjects revealed that the amplitude and the latency of the N1m exhibits directional tuning to sound location, with the amplitude of the right-hemispheric N1m being particularly sensitive to the amount of spatial cues in the stimuli. The BAR, HRTF, and combined ITD + ILD stimuli resulted both in a larger dynamic range and in a more systematic distribution of the N1m amplitude across stimulus angle than did the ITD or ILD stimuli alone. Further, the right-hemispheric source loci of the N1m responses for the BAR and HRTF stimuli were anterior to those for the ITD and ILD stimuli. In behavioral tests, we measured the ability of the subjects to localize BAR and HRTF stimuli in terms of azimuthal error and front-back confusions. We found that behavioral performance correlated positively with the amplitude of the N1m. Thus, the activity taking place already in the auditory cortex predicts behavioral sound detection of spatial stimuli, and the amount of spatial cues embedded in the signal are reflected in the activity of this brain area.

Acoustic Stimulation↗

The role of blind humans' visual cortex in auditory change detection.

Several studies using brain imaging have demonstrated occipital-cortex activation in blind individuals during tactile and auditory tasks, suggesting that the visual cortex deprived of its normal input has adopted a new role in information processing. So far, however, at what stages of information processing and to which perceptual sub-processes this applies remains unclear. We determined the auditory functions of this cortical region in early-blind humans by means of functional magnetic resonance imaging. We found that these areas were not activated by the mere presence of sound, but were involved in the attentive processing of changes in the auditory environment, which is important in detecting potentially dangerous or other important events in the surroundings, for example.

Acoustic Stimulation↗

Magnetic fields evoked by speech sounds in preschool children.

OBJECTIVE: Our objective was to study how well the auditory evoked magnetic fields (EF) reflect the behavioral discrimination of speech sounds in preschool children, and if they reveal the same information as simultaneously recorded evoked potentials (EP). METHODS: EFs and EPs were recorded in 11 preschool children (mean age 6 years 9 months) using an oddball paradigm with two sets of speech stimuli consisting both of one standard and two deviants. After the brain activity recording, children were tested on behavioural discrimination of the same stimuli presented in pairs. RESULTS: There was a mismatch negativity (MMN) calculated from difference curves and its magnetic counterpart MMNm measured from the original responses only to those deviants, which were behaviourally easiest to discriminate from the standards. In addition, EF revealed significant differences between the locations of the activation depending on the hemisphere and stimulus properties. CONCLUSIONS: EF, in addition to reflecting the sound-discrimination accuracy in a similar manner as EP, also reflected the spatial differences in activation of the temporal lobes. SIGNIFICANCE: These results suggest that both EPs and EFs are feasible for investigating the neural basis of sound discrimination in young children. The recording of EFs with its high spatial resolution reveals information on the location of the activated neural sources.

Acoustic Stimulation↗