PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Speech Recognition Software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Use of speech recognition software: a vocal endurance test for the new millennium?

Speech recognition software for the personal or office computer is a relatively new area of technology. As the number of these products has increased so has use of this software. Some individuals will employ speech recognition systems due to difficulty with the conventional keyboard and mouse interface: others will use it for perceived efficiency or simply novelty. Regardless of the reason for use of this technology, the voice demands associated with extended or frequent use can be high, placing the user at risk for vocal difficulties. This paper reviews the case of an individual referred to our multidisciplinary voice care program for evaluation and treatment of vocal difficulties that began secondary to utilization of speech recognition software. We discuss medical and vocal histories, examination findings, treatment, and treatment outcomes.

Adult↗

Speech recognition software.

This article discusses the use of speech recognition software by means of reviewing two leading packages. Both programs require considerable training before they can be used effectively, but are then able to convert continuous speech into text with varying degrees of success.

CD-ROM↗

Comparative evaluation of three continuous speech recognition software packages in the generation of medical reports.

OBJECTIVE: To compare out-of-box performance of three commercially available continuous speech recognition software packages: IBM ViaVoice 98 with General Medicine Vocabulary; Dragon Systems NaturallySpeaking Medical Suite, version 3.0; and L&H Voice Xpress for Medicine, General Medicine Edition, version 1.2. DESIGN: Twelve physicians completed minimal training with each software package and then dictated a medical progress note and discharge summary drawn from actual records. MEASUREMENTS: Errors in recognition of medical vocabulary, medical abbreviations, and general English vocabulary were compared across packages using a rigorous, standardized approach to scoring. RESULTS: The IBM software was found to have the lowest mean error rate for vocabulary recognition (7.0 to 9.1 percent) followed by the L&H software (13.4 to 15.1 percent) and then Dragon software (14.1 to 15.2 percent). The IBM software was found to perform better than both the Dragon and the L&H software in the recognition of general English vocabulary and medical abbreviations. CONCLUSION: This study is one of a few attempts at a robust evaluation of the performance of continuous speech recognition software. Results of this study suggest that with minimal training, the IBM software outperforms the other products in the domain of general medicine; however, results may vary with domain. Additional training is likely to improve the out-of-box performance of all three products. Although the IBM software was found to have the lowest overall error rate, successive generations of speech recognition software are likely to surpass the accuracy rates found in this investigation.

Evaluation Studies as Topic↗

Free-text data entry by speech recognition software and its impact on clinical routine.

We conducted a study to evaluate speech recognition software in an otorhinolaryngology unit and to assess its impact on productivity prior to general implementation. Current speech recognition software (IBM ViaVoice, version 10) was implemented on a personal computer with a 2-GHz central processing unit, 256 MB of RAM, and a 30-GB hard disk drive, with and without add-on professional vocabulary for otorhinolaryngology. This vocabulary was added by the automated analysis of an additional 12,257 documents from our department. We compared the word recognition error rates for three different text types and determined their impact on the amount of surgeon's time that was invested in the production of an error-free document. Although error rates without any professional vocabulary database were rather high (operation reports: 38.72%; consultation notes: 27.77%), the patient information was edited with a satisfactory result (10.65%). Best results were obtained with the specialty-related vocabulary database added by the analysis of our own documents (operation reports: 5.45%; consultation notes: 5.21%). An increase in productivity compared with that of conventional transcription was found at an error rate of less than 16%.

Efficiency, Organizational↗

Speech recognition software as an assistive device: a pilot study of user satisfaction and psychosocial impact.

The purpose of this study was to gather data concerning the psychosocial (quality of life) impact of speech recognition software on individuals with physical disabilities and to identify how satisfied these individuals were with this software as a computer access method. Two standardized questionnaires, the Psychosocial Impact of Assistive Devices Scale (PIADS) and the Quebec User Evaluation of Satisfaction with assistive Technology (QUEST) were administered to ten participants with physical disabilities who received speech recognition software following an assistive technology evaluation. The results of this study indicated that 90% of the participants were quite satisfied with speech recognition software as an assistive device and that the software had a somewhat positive psychosocial impact on their lives. Four themes emerged concerning what the participants liked most about the software: 1) the software provided a method of access when they were not previously accessing a computer, 2) the software increased independence, 3) the software made computer use more efficient, and 4) the software provided a choice or flexibility in computer access. Although this study demonstrated that these speech recognition software users are generally satisfied with the software and it has had a positive impact on their life, it also suggests that there is a need to examine the role of training on satisfaction and successful use of the software.

Adult↗

Combining speech recognition software with Digital Imaging and Communications in Medicine (DICOM) workstation software on a Microsoft Windows platform.

This presentation describes our experience in combining speech recognition software, clinical review software, and other software products on a single computer. Different processor speeds, random access memory (RAM), and computer costs were evaluated. We found that combining continuous speech recognition software with Digital Imaging and Communications in Medicine (DICOM) workstation software on the same platform is feasible and can lead to substantial savings of hardware cost. This combination optimizes use of limited workspace and can improve radiology workflow.

Humans↗

Speech recognition technology: an outlook for human-to-machine interaction.

Speech recognition, as an enabling technology in healthcare-systems computing, is a topic that has been discussed for quite some time, but is just now coming to fruition. Traditionally, speech-recognition software has been constrained by hardware, but improved processors and increased memory capacities are starting to remove some of these limitations. With these barriers removed, companies that create software for the healthcare setting have the opportunity to write more successful applications. Among the criticisms of speech-recognition applications are the high rates of error and steep training curves. However, even in the face of such negative perceptions, there remains significant opportunities for speech recognition to allow healthcare providers and, more specifically, physicians, to work more efficiently and ultimately spend more time with their patients and less time completing necessary documentation. This article will identify opportunities for inclusion of speech-recognition technology in the healthcare setting and examine major categories of speech-recognition software--continuous speech recognition, command and control, and text-to-speech. We will discuss the advantages and disadvantages of each area, the limitations of the software today, and how future trends might affect them.

Health Care Sector↗

[Design and implementation of aphasia rehabilitation software based on speech recognition].

A new software therapy instrument is proposed for the rehabilitation of aphasia caused by cerebral disorder, which is different from general drug therapy or physical therapy and is, based on modern speech and biofeedback principles. Aphasia rehabilitation software package on the therapy instrument were designed and implement. The features of the software and its future application were discussed.

Aphasia↗

The effect of speech recognition on working postures, productivity and the perception of user friendliness.

A comparative, experimental study with repeated measures has been conducted to evaluate the effect of the use of speech recognition on working postures, productivity and the perception of user friendliness. Fifteen subjects performed a standardised task, first with keyboard and mouse and, after a six week training period, with speech recognition. The use of speech recognition leads to improved postures of wrist, forearm, upper arm and shoulder and improvement of neck movements when compared to the use of keyboard and mouse. Although the observation method was basic, this study provides insight into the potential benefits speech recognition has for posture. However, productivity decreased for most subjects and speech recognition appears to be usable for specific tasks only. From the perspective of productivity and the perception of user friendliness further development of speech recognition software is necessary. Up to now, speech recognition seems especially beneficial for people with WMSD complaints.

Adult↗

Laboratory voice data entry system.

We have assembled a system using a personal computer workstation equipped with standard office software, an audio system, speech recognition software and an inexpensive radio-based wireless microphone that permits laboratory workers to enter or modify data while performing other work. Speech recognition permits users to enter data while their hands are holding equipment or they are otherwise unable to operate a keyboard. The wireless microphone allows unencumbered movement around the laboratory without a "tether" that might interfere with equipment or experimental procedures. To evaluate the potential of voice data entry in a laboratory environment, we developed a prototype relational database that records the disposal of radionuclides and/or hazardous chemicals. Current regulations in our laboratory require that each such item being discarded must be inventoried and documents must be prepared that summarize the contents of each container used for disposal. Using voice commands, the user enters items into the database as each is discarded. Subsequently, the program prepares the required documentation.

Computers↗

Computerized content analysis of speech plus speech recognition in the measurement of neuropsychiatric dimensions.

The Psychiatric Content Analysis and Diagnosis (PCAD) program performs automated content analysis of machine-readable transcriptions of speech samples to measure the magnitude of neuropsychiatric states and traits. Technological advances provided by computerized speech recognition may offer a possible alternative to labor-intensive manual transcription for preparation of samples for PCAD processing. To test this hypothesis, 25 digitally recorded verbal samples were transcribed both manually and by a commercially available speech recognition software package, and the transcriptions scored by PCAD. The inter-correlations between scores derived from the two different methods of transcriptions offer mixed results, with values ranging from a high of 0.920 to a low of -0.119.

Cognition Disorders↗

Speech recognition interface to a hospital information system using a self-designed visual basic program: initial experience.

Speech recognition (SR) in the radiology department setting is viewed as a method of decreasing overhead expenses by reducing or eliminating transcription services and improving care by reducing report turnaround times incurred by transcription backlogs. The purpose of this study was to show the ability to integrate off-the-shelf speech recognition software into a Hospital Information System in 3 types of military medical facilities using the Windows programming language Visual Basic 6.0 (Microsoft, Redmond, WA). Report turnaround times and costs were calculated for a medium-sized medical teaching facility, a medium-sized nonteaching facility, and a medical clinic. Results of speech recognition versus contract transcription services were assessed between July and December, 2000. In the teaching facility, 2042 reports were dictated on 2 computers equipped with the speech recognition program, saving a total of US dollars 3319 in transcription costs. Turnaround times were calculated for 4 first-year radiology residents in 4 imaging categories. Despite requiring 2 separate electronic signatures, we achieved an average reduction in turnaround time from 15.7 hours to 4.7 hours. In the nonteaching facility, 26600 reports were dictated with average turnaround time improving from 89 hours for transcription to 19 hours for speech recognition saving US dollars 45500 over the same 6 months. The medical clinic generated 5109 reports for a cost savings of US dollars 10650. Total cost to implement this speech recognition was approximately US dollars 3000 per workstation, mostly for hardware. It is possible to design and implement an affordable speech recognition system without a large-scale expensive commercial solution.

Cost-Benefit Analysis↗

Muscle tension dysphonia in patients who use computerized speech recognition systems.

The use of speech recognition systems as a replacement for other types of transcription systems is increasing rapidly, partly because many people are unable to use conventional keyboards as a result of upper-extremity repetitive strain injury (RSI). However, the frequent or continuous use of such systems can cause muscle tension dysphonia in some patients. The scientific literature suggests that there is an association between upper-extremity RSI and muscle tension dysphonia. We present a retrospective case series of five patients with workplace upper-extremity RSI who developed muscle tension dysphonia soon after they began using discrete computerized speech recognition software. The diagnosis of dysphonia was based on laryngovideostroboscopy, acoustic analyses, and voice load testing. All patients had normal voice when using everyday speech, but speaking into the computer resulted in the rapid onset of aperiodicity, strain, and a decrease in fundamental frequency. In three of the five patients, laryngovideostroboscopy showed posterior glottic overapproximation, but no other abnormalities. Treatment was centered on voice therapy and avoidance of long periods of using computerized speech recognition systems. The condition of three of the five patients improved with therapy. We conclude that computer speech recognition programs can lead to the onset of muscle tension dysphonia in some patients. These patients can be successfully treated with voice therapy.

Adult↗

Automatic concept extraction from spoken medical reports.

OBJECTIVE: The objective of this project is to investigate methods whereby a combination of speech recognition and automated indexing methods substitute for current transcription and indexing practices. METHODS: We based our study on existing speech recognition software programs and on NOMINDEX, a tool that extracts MeSH concepts from medical text in natural language and that is mainly based on a French medical lexicon and on the UMLS. For each document, the process consists of three steps: (1) dictation and digital audio recording, (2) speech recognition, (3) automatic indexing. The evaluation consisted of a comparison between the set of concepts extracted by NOMINDEX after the speech recognition phase and the set of keywords manually extracted from the initial document. The method was evaluated on a set of 28 patient discharge summaries extracted from the MENELAS corpus in French, corresponding to in-patients admitted for coronarography. RESULTS: The overall precision was 73% and the overall recall was 90%. Indexing errors were mainly due to word sense ambiguity and abbreviations. A specific issue was the fact that the standard French translation of MeSH terms lacks diacritics. A preliminary evaluation of speech recognition tools showed that the rate of accurate recognition was higher than 98%. Only 3% of the indexing errors were generated by inadequate speech recognition. DISCUSSION: We discuss several areas to focus on to improve this prototype. However, the very low rate of indexing errors due to speech recognition errors highlights the potential benefits of combining speech recognition techniques and automatic indexing.

Abstracting and Indexing↗

Experimental analysis of human vocal behavior: applications of speech-recognition technology.

Recent developments in speech recognition make it feasible to apply the technology to study vocal behavior. The present study illustrates the use of this technology to establish functional stimulus classes. Eight students were taught to say nonsense words in the presence of arbitrarily assigned sets of symbols consistent with three three-member experimenter-defined stimulus classes. Computer-controlled speech-recognition software was used to record, analyze, and differentially reinforce vocal responses. When the stimulus classes were established, students were taught to say a new nonsense word in the presence of one member of each stimulus class. Transfer of function was tested subsequently to determine if the novel stimulus names transferred to the remaining stimulus class members. Most subjects required two iterations of the training and testing procedures before transfer occurred. The data illustrate the usefulness of recording vocal behavior during stimulus control procedures and demonstrate the use of speech-recognition technology. The paper also describes the current state of speech-recognition technology and suggests several other areas of research that might benefit from using vocal behavior as its primary datum.

Adolescent↗