PubMed Health⌕ Search

Biomedical subjects

J Kreiman

Publications and source records attributed to J Kreiman.

At least 19 recordsLinked to original sources

Measuring vocal quality with speech synthesis.

Much previous research has demonstrated that listeners do not agree well when using traditional rating scales to measure pathological voice quality. Although these findings may indicate that listeners are inherently unable to agree in their perception of such complex auditory stimuli, another explanation implicates the particular measurement method-rating scale judgments-as the culprit. An alternative method of assessing quality-listener-mediated analysis-synthesis-was devised to assess this possibility. In this new approach, listeners explicitly compare synthetic and natural voice samples, and adjust speech synthesizer parameters to create auditory matches to voice stimuli. This method is designed to replace unstable internal standards for qualities like breathiness and roughness with externally presented stimuli, to overcome major hypothetical sources of disagreement in rating scale judgments. In a preliminary test of the reliability of this method, listeners were asked to adjust the signal-to-noise ratio for 12 synthetic pathological voices so that the resulting stimuli matched the natural target voices as well as possible For comparison to the synthesis judgments, listeners also judged the noisiness of the natural stimuli in a separate task using a traditional visual-analog rating scale. For 9 of the 12 voices, agreement among listeners was significantly (and substantially) greater for the synthesis task than for the rating scale task. Response variances for the two tasks did not differ for the remaining three voices. However, a second experiment showed that the synthesis settings that listeners selected for these three voices were within a difference limen, and therefore observed differences were perceptually insignificant. These results indicate that listeners can in fact agree in their perceptual assessments of voice quality, and that analysis-synthesis can measure perception reliably.

Adult↗

Sources of listener disagreement in voice quality assessment.

Traditional interval or ordinal rating scale protocols appear to be poorly suited to measuring vocal quality. To investigate why this might be so, listeners were asked to classify pathological voices as having or not having different voice qualities. It was reasoned that this simple task would allow listeners to focus on the kind of quality a voice had, rather than how much of a quality it possessed, and thus might provide evidence for the validity of traditional vocal qualities. In experiment 1, listeners judged whether natural pathological voice samples were or were not primarily breathy and rough. Listener agreement in both tasks was above chance, but listeners agreed poorly that individual voices belonged in particular perceptual classes. To determine whether these results reflect listeners' difficulty agreeing about single perceptual attributes of complex stimuli, listeners in experiment 2 classified natural pathological voices and synthetic stimuli (varying in f0 only) as low pitched or not low pitched. If disagreements derive from difficulties dividing an auditory continuum consistently, then patterns of agreement should be similar for both kinds of stimuli. In fact, listener agreement was significantly better for the synthetic stimuli than for the natural voices. Difficulty isolating single perceptual dimensions of complex stimuli thus appears to be one reason why traditional unidimensional rating protocols are unsuited to measuring pathologic voice quality. Listeners did agree that a few aphonic voices were breathy, and that a few voices with prominent vocal fry and/or interharmonics were rough. These few cases of agreement may have occurred because the acoustic characteristics of the voices in question corresponded to the limiting case of the quality being judged. Values of f0 that generated listener agreement in experiment 2 were more extreme for natural than for synthetic stimuli, consistent with this interpretation.

Adolescent↗

Exit jet particle velocity in the in vivo canine laryngeal model with variable nerve stimulation.

This study extends previous work on exit jet particle velocity in the in vivo canine model of phonation by measuring air particle velocity at multiple locations in the midline of the glottis and across multiple levels of recurrent laryngeal nerve (RLN) and superior laryngeal nerve (SLN) stimulation. In a second experiment, exit jet particle velocity was measured at midline and offmidline positions with constant levels of RLN and SLN stimulation. In this study, peak particle velocity was higher at the anterior commissure than at the posterior commissure in the midline of the glottis, and peak particle velocity was higher at the midline than at offmidline positions. In addition, increasing levels of RLN stimulation resulted in increasing peak particle velocity; however, increasing levels of SLN stimulation failed to produce a uniform effect on peak particle velocity.

Air↗

Treatment of Parkinson hypophonia with percutaneous collagen augmentation.

OBJECTIVES: It has been estimated that more than 70% of patients with Parkinson disease experience voice and speech disorders characterized by weak and breathy phonation, and dysarthria. This study reports on the efficacy of treating Parkinson patients who have glottal insufficiency. STUDY DESIGN AND METHODS: Thirty-five patients underwent collagen augmentation of the vocal folds for hypophonia associated with Parkinson disease, using a new technique of percutaneous injection with fiberoptic guidance. Patient response to the collagen augmentation was determined by telephone survey. RESULTS AND CONCLUSIONS: The procedure required minimal patient participation and was safely performed on all the patients who were studied. Results of the survey indicated that 75% of patient responses demonstrated satisfaction with the technique, compared with 16% of patient ratings reflecting dissatisfaction. These results were moderately correlated with the duration of improvement of the dysphonia. Results of this preliminary evaluation demonstrate that voice deficits in Parkinson disease are amenable to vocal fold augmentation. Because this procedure requires minimal patient participation and can be safely performed in an office setting, it may also be useful in other severely debilitating neuromotor diseases that result in glottal insufficiency and hypophonia.

Aged↗

Combined arytenoid adduction and laryngeal reinnervation in the treatment of vocal fold paralysis.

OBJECTIVE/HYPOTHESIS: Glottal closure and symmetrical thyroarytenoid stiffness are two important functional characteristics of normal phonatory posture. In the treatment of unilateral vocal cord paralysis, vocal fold medialization improves closure, facilitating entrainment of both vocal folds for improved phonation, and reinnervation is purported to maintain vocal fold bulk and stiffness. A combination of medialization and reinnervation would be expected to further improve vocal quality over medialization alone. STUDY DESIGN: A retrospective review of preoperative and postoperative voice analysis on all patients who underwent arytenoid adduction alone (adduction group) or combined arytenoid adduction and ansa cervicalis to recurrent laryngeal nerve anastomosis (combined group) between 1989 and 1995 for the treatment of unilateral vocal cord paralysis. Patients without postoperative voice analysis were invited back for its completion. A perceptual analysis was designed and completed. METHODS: Videostroboscopic measures of glottal closure, mucosal wave, and symmetry were rated. Aerodynamic parameters of laryngeal airflow and subglottic pressure were measured. A 2-second segment of sustained vowel was used for perceptual analysis by means of a panel of voice professionals and a rating system. Statistical calculations were performed at a significance level of P = .05. RESULTS: There were 9 patients in the adduction group and 10 patients in the combined group. Closure and mucosal wave improved significantly in both groups. Airflow decreased in both groups, but the decrease reached statistical significance only in the adduction group. Subglottic pressure remained unchanged in both groups. Both groups had significant perceptual improvement of voice quality. In all tested parameters the extent of improvement was similar in both groups. CONCLUSION: The role of laryngeal reinnervation in the treatment of unilateral vocal cord paralysis remains to be established.

Adult↗

Validity of rating scale measures of voice quality.

The validity of perceptual measures of vocal quality has been neglected in studies of voice, which focus more commonly on rater reliability. Validity depends in part on reliability, because an unreliable test does not measure what it is intended to measure. However, traditional measures of rating reliability only partially represent interrater agreement, because they cannot reflect variations or patterns of agreement for specific voice samples. In this paper the likelihood that two raters would agree in their ratings of a single voice is examined, for each voice in five previously gathered data sets. Results do not support the continued assumption that traditional rating procedures produce useful indices of listeners' perceptions. Listeners agreed very poorly in the midrange of scales for breathiness and roughness, and mean ratings in the midrange of such scales did not represent the extent to which a voice possesses a quality, but served only to indicate that listeners disagreed. Techniques like analysis by synthesis or judgment of similarity avoid decomposing quality into constituent dimensions, and do not require a listener to compare an external stimulus to an unstable internal representation, thus decreasing the error in measures of quality. Modeling individual differences in perception can increase the variance accounted for in models of quality, further reducing the error in perceptual measures. Thus such techniques may provide valid alternatives to current approaches.

Humans↗

Functional differences between the two bellies of the cricothyroid muscle.

The contraction of the cricothyroid (CT) muscle, which results in a decrease in the distance between the thyroid and cricoid cartilages, is considered to be the main factor in lengthening the vocal folds. This is achieved by rotation of the CT joint. The CT muscle is composed of two distinct bellies, the pars recta and the pars obliqua. The function of each subunit is not clearly understood, although it is believed that they act differently because their fibers run in different directions. To clarify the function of the two bellies in phonation, the fundamental frequency (F0), vocal intensity, subglottic pressure, vocal fold length, and CT distance were measured using an in vivo canine laryngeal model. On the basis of these measurements, we demonstrated that the two bellies are varied in their effect on raising the pitch, rotation, and forward translation of the CT joint. The stimulation of the pars recta nerve resulted in a greater increase in the F0 value compared with that of pars obliqua. The combined activity of the pars recta and pars obliqua is important in adjustment of the vocal fold length. The CT approximations directed parallel to the pars recta and pars obliqua simultaneously were more effective in elevation of the pitch than the approximation placed parallel to the pars recta only. This finding may be clinically significant with regard to CT approximation thyroplasty in human trails.

Analysis of Variance↗

Vocal function following vertical hemilaryngectomy: comparison of four reconstruction techniques in the canine.

The goals of laryngeal reconstruction have been prevention of aspiration, production of a functional voice, and maintenance of an adequate airway for decannulation. A number of procedures for partial laryngeal reconstruction have accomplished these objectives. However, few studies have attempted to compare patients' vocal characteristics following different reconstruction procedures. In this study, an in vivo canine model was used to compare acoustic and aerodynamic measures of vocal function for the following vertical hemilaryngectomy reconstruction techniques: 1) a superiorly based sternohyoid muscle flap, 2) a modified epiglottic laryngoplasty, 3) a new procedure using a layered vascularized buccal mucosal flap and a transversely oriented sternohyoid muscle flap, and 4) hemilaryngeal transplantation combined with arytenoid adduction. Hemitransplantation provided the most efficient phonation of the four techniques. The vascularized buccal mucosa flap produced the best phonation of the autologous tissue techniques examined. Both vascularized buccal mucosa flap and hemilaryngeal transplantation subjects demonstrated a mucosal wave on stroboscopy. The results indicate that vocal function will improve as the layered structure of the vocal fold is more accurately replicated in a reconstructed hemilarynx. Endoscopic findings and whole organ sections are presented.

Animals↗

Effects of driving pressure and recurrent laryngeal nerve stimulation on glottic vibration in a constant pressure model.

Glottic phonatory parameters have been studied in constant flow models; however, the lung-thorax system is better viewed as a constant pressure source. Adjusting the driving pressure and recurrent laryngeal nerve stimulation as independent variables, rather than as dependent variables, may provide a more physiologic understanding of laryngeal function and glottic parameters, including subglottic pressure, airflow, fundamental frequency, and glottic area. In three dogs subglottic pressure and airflow were measured in two separate conditions: with constant recurrent laryngeal nerve stimulation and varying driving pressure, and with constant driving pressure and varying recurrent laryngeal nerve stimulation. Videostroboscopic measures on four dogs assessed glottic areas with constant recurrent laryngeal nerve stimulation at different driving pressures. With constant recurrent laryngeal nerve stimulation, increasing driving pressure had no effect on glottic areas, whereas subglottic pressure, fundamental frequency, and airflow increased significantly. However, changes in subglottic pressure were minimal in comparison with changes in driving pressure. At constant driving pressure, increasing recurrent laryngeal nerve stimulation increased subglottic pressure and fundamental frequency and decreased airflow. These findings suggest that during phonation subglottic pressure is primarily dependent on recurrent laryngeal nerve stimulation and laryngeal muscular contraction, but not on lung driving pressure.

Analysis of Variance↗

Comparison of voice analysis systems for perturbation measurement.

Dysphonic voices are often analyzed using automated voice analysis software. However, the reliability of acoustic measures obtained from these programs remains unknown, particularly when they are applied to pathological voices. This study compared perturbation measures from CSpeech, Computerized Speech Laboratory, SoundScope, and a hand marking voice analysis system. Sustained vowels from 29 male and 21 female speakers with mild to severe dysphonia were digitized, and fundamental frequency (F0), jitter, shimmer, and harmonics- or signal-to-noise ratios were computed. Commercially available acoustical analysis programs agreed well, but not perfectly, in their measures of F0. Measures of perturbation in the various analysis packages use different algorithms, provide results in different units, and often yield values for voices that violate the assumption of quasi-periodicity. As a result, poor rank order correlations between programs using similar measures of perturbation were noted. Because measures of aperiodicity apparently cannot be reliably applied to voices that are even mildly aperiodic, we question their utility in quantifying vocal quality, especially in pathological voices.

Female↗

Characteristics of an in vivo canine model of phonation with a constant air pressure source.

Many previous studies of laryngeal biomechanics using in vivo models have employed a constant air How source. Several authors have recently suggested that the lung-thorax system functions as a constant pressure source during phonation. This study describes an in vivo canine system designed to maintain a constant peak subglottic pressure (Psub) using a pressure-controlling mechanism. Increasing levels of recurrent laryngeal nerve (RLN) stimulation resulted in a significant rise in resistance followed by a plateau. For a given Psub, flow decreased significantly and precipitously with increasing stimulation and then quickly plateaued. Vocal intensity increased with increasing RLN stimulation until a peak was reached. After this peak, intensity dropped until a plateau was reached, corresponding to the flow minimum. At a given Psub, increasing levels of RLN stimulation resulted in a normal distribution of vocal efficiencies.

Airway Resistance↗

The perceptual structure of pathologic voice quality.

Although perceptual assessment is included in most protocols for evaluating pathologic voices, a standard set of valid scales for measuring voice quality has never been established. Standardization is important for theory and for clinical acceptance, and also because validation of objective measures of voice depends on valid perceptual measures. The present study used large sets (n = 80) of male and female voices, representing a broad range of diagnoses and vocal severities. Eight experts judged the dissimilarity of each pair of voices, and responses were analyzed using nonmetric individual differences multidimensional scaling. Results indicate that differences between listeners in perceptual strategy are so great that the fundamental assumption of a common perceptual space must be questioned. Because standardization depends on the assumption that listeners are similar, it is concluded that efforts to standardize perceptual labels for voice quality are unlikely to succeed. However, analysis by synthesis may provide an alternate means of modeling quality as a function of both voices and listeners, thus avoiding this problem.

Adult↗

Comparing reliability of perceptual ratings of roughness and acoustic measure of jitter.

Acoustic analysis is often favored over perceptual evaluation of voice because it is considered objective, and thus reliable. However, recent studies suggest this traditional bias is unwarranted. This study examined the relative reliability of human listeners and automatic systems for measuring perturbation in the evaluation of pathologic voices. Ten experienced listeners rated the roughness of 50 voice samples (ranging from normal to severely disordered) on a 75 mm visual analog scale. Rating reliability within and across listeners was compared to the reliability of jitter measures produced by several voice analysis systems (CSpeech, SoundScope, CSL, and an interactive hand-marking system). Results showed that overall listeners agreed as well or better than "objective" algorithms. Further, listeners disagreed in predictable ways, whereas automatic algorithms differed in seemingly random fashions. Finally, listener reliability increased with severity of pathology; objective methods quickly broke down as severity increased. These findings suggest that listeners and analysis packages differ greatly in their measurement characteristics. Acoustic measures may have advantages over perceptual measures for discriminating among essentially normal voices; however, reliability is not a good reason for preferring acoustic measures of perturbation to perceptual measures.

Female↗

The effect of gas density on glottal vibration and exit jet particle velocity.

Although theoretical studies include a term for gas density in their mathematical descriptions of glottal aerodynamics, the effect of gas density on glottal vibration has not been examined empirically. In this study, an in vivo canine model was used to evaluate the effect of gas density on glottal vibration by comparing phonation with air and helium. With gas flow and nerve stimulation held constant, phonation with helium resulted in an increased exit jet particle velocity for helium (45 m/s) compared to air (34 m/s). However, the measured increase in helium velocity was less than predicted by a proportional relationship between transglottal pressure and dynamic pressure. This difference could be due to a change in the constant of proportionality or in the dynamic pressure loss coefficient associated with the use of helium.

Animals↗

Relation of recurrent laryngeal nerve compound action potential to laryngeal biomechanics.

This study was designed to investigate the compound action potential (CAP) of the recurrent laryngeal nerve (RLN) and to correlate this electrophysiologic signal to laryngeal biomechanics and phonatory function. Four adult mongrel canines were anesthetized. The RLN was isolated and stimulated, and recording electrodes were applied. The electromyographic (EMG) electrode was placed in the thyroarytenoid (TA) muscle. The RLN CAP and the EMG of the TA muscle were recorded and compared to the stimulation intensity, subglottic pressure (Psub), and each other. The CAP peak-to-peak and EMG peak-to-peak amplitudes demonstrated a sigmoidal relation to stimulus intensity and a linear relation to Psub and to each other. On the basis of these findings, the RLN CAP appears to be a reliable physiologic measure of laryngeal function.

Action Potentials↗

The multidimensional nature of pathologic vocal quality.

Although the terms "breathy" and "rough" are frequently applied to pathological voices, widely accepted definitions are not available and the relationship between these qualities is not understood. To investigate these matters, expert listeners judged the dissimilarity of pathological voices with respect to breathiness and roughness. A second group of listeners rated the voices on unidimensional scales for the same qualities. Multidimensional scaling analyses suggested that breathiness and roughness are related, multidimensional constructs. Unidimensional ratings of both breathiness and roughness were necessary to describe patterns of similarity with respect to either quality. Listeners differed in the relative importance given to different aspects of voice quality, particularly when judging roughness. The presence of roughness in a voice did not appear to influence raters' judgments of breathiness; however, judgments of roughness were heavily influenced by the degree of breathiness, the particular nature of the influence varying from listener to listener. Differences in how listeners focus their attention on the different aspects of multidimensional perceptual qualities apparently are a significant source of interrater unreliability (noise) in voice quality ratings.

Humans↗

Changes in glottal area associated with increasing airflow.

Laryngeal resistance varies inversely with airflow during phonation. This study evaluated the morphological changes in the glottis that accompany decreases in laryngeal resistance at higher levels of airflow. An in vivo canine model of phonation and a video analysis system were used to assess changes in area. Four animals were examined stroboscopically as airflow increased, with constant recurrent laryngeal nerve stimulation. Glottal dynamics were evaluated by means of photoglottography, electroglottography, and measures of subglottic pressure. Analysis of digitized stroboscopic images indicated that increasing airflow had no obvious effect on the glottal chink (vocal process contact). Increasing airflow was associated with an increase in the area of peak opening and an increase in the glottal area integral.

Analysis of Variance↗

Function of the interarytenoid muscle in a canine laryngeal model.

The interarytenoid (IA) muscle has rarely been studied in the living larynx. In this work, the role of the IA muscle in phonation was studied in three dogs by means of an in vivo phonation model. The isolated action of the IA muscle was studied by sectioning and stimulating its nerve branch. As IA activity increased, subglottic pressure increased significantly until a plateau was reached. In the absence of superior laryngeal nerve stimulation, the fundamental frequency rose with increasing IA activity. In the presence of superior laryngeal nerve stimulation, however, no significant change in fundamental frequency was observed with increasing IA activity. Measurement of adductory force demonstrated that the IA muscle adducts primarily the posterior vocal fold. In this canine model, phonation was not possible without IA stimulation, owing to a large posterior glottic chink.

Analysis of Variance↗