PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Software Validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Toward an intelligent wound assessment system.

There is general agreement regarding the need for pressure ulcer assessment methodology which more discretely reflects relevant aspects of wound status than does the commonly used staging system. The Pressure Sore Status Tool (PSST) is one such instrument which was developed with consensual expert input. While the psychometric properties of the PSST have been reported in the literature, the instrument was validated using ET nurses, highly trained wound care specialists, and existed only in manual form. This paper reports results from attempts to establish reliability estimates for healthcare practitioners without extraordinary wound care training or experience. The paper further describes the automation of the PSST and provides examples of pressure ulcer profiles tracked over time. Results indicate that inter-rater reliability with general healthcare practitioners was .78 and intra-rater reliability was .89. The practitioners were able to use the PSST for over six months and the automated system allowed analysis of wound healing profiles that would have been difficult using a manual system. These results imply that movement toward an automated system which makes discriminations regarding the effects of various treatment and intervention strategies is possible and practical.

Aged↗

The role of quality assurance in computer inspections.

Changing technology affords the Quality Assurance auditor with the challenge of applying computer validation concepts to a variety of computer system types. In addition, these technology changes have caused the developers role to change as well. In an innovative research facility, the developer may include an in-house professional group, a vendor, or an end-user. With these issues in mind, the QA auditor needs a tool to accomplish the task of inspecting systems as they are being created. The prospective inspection process is the tool for accomplishing this task. This inspection involves the QA auditor's involvement in the development of a new system, as a member of the development team, from the initial creation through the implementation of the computer system. This presentation will focus on illustrating the steps in conducting the prospective inspection process, from expected deliverables and document reviews to final report and management notification. The benefits of QA involvement in the development process will also be discussed.

Clinical Laboratory Information Systems↗

The performance of the knowledge-based system VALAB revisited: an evaluation after five years.

In 1988, inundated by the tedious work of validation of laboratory reports in a large hospital biochemistry laboratory, we designed VALAB, a knowledge-based system specially dedicated to this iterative function. Coping at first with a few biochemical tests, the program has been progressively expanded to forty-five common chemical tests. Simultaneously some new rules have been introduced to "weight" the conclusion in different circumstances and rules taking into consideration some clinical data have also been written. Moreover the program moved to other disciplines, pH and blood gases, haematology and coagulation. Accordingly the evaluation protocol has been modified, incorporating a new step, the consensus decision of the pathologists, operating within the initial protocol and based upon the various criteria of epidemiology. These major changes and improvements have led us to check and describe again the performance of this updated VALAB knowledge-based system.

Artificial Intelligence↗

Pattern recognition in health insurance claims databases.

Information in claims databases resides in data patterns rather than in data elements. Finding this information requires new terminology, a willingness to pose questions of form rather than specific hypotheses, and a quality control system that elevates the correctness of data relations above the validity of single facts. The language of claims data is a newspeak of CPT (Current Procedural Terminology), HCPCS (Health Care Financing Agency Common Procedure Coding System), ICD (International Classification of Disease), and NDC (National Drug Codes) for pharmaceutical codes. The techniques of pattern discovery are really ways of asking the data for classes of relations, and they vary in their reliance on external information. Sometimes, the question is entirely constrained by preceding factors. Other times we may recast the natural history of disease into a claims context and ask the data to give us the shape of disease evolution. We can use highly automated systems to evaluate the relations between prespecified factors, or empirical techniques to search out common relations that we have not specified in advance. Using massive data sets requires that quality control corresponds to the nature of the high-level information that we derive from large databases.

Databases as Topic↗

Computerized decision support for concurrent utilization review using the HELP system.

OBJECTIVE: Development and evaluation of computerized concurrent utilization review (UR) support taking advantage of a clinically rich computerized patient database. DESIGN: The Automated Support System for Utilization Review (ASSURE) applies the Appropriateness Evaluation Protocol (AEP) Day of Care criteria to computerized patient data in the HELP hospital information system. This paper reports the development, verification, and validation of ASSURE. MEASUREMENTS: Implementation correctness was verified by measuring agreement with a nurse reviewer, using separate sample sets for all 20 criteria for a total of 560 current inpatients. Usefulness in detecting inappropriate days of care was validated by two nurse reviewers who were crossed with manual and computer-assisted review methods in a blocked design for 168 current inpatients. Agreement with reviewers, sensitivity, specificity, positive predictive value, and negative predictive value were measured. RESULTS: Agreement was very good for satisfaction of criteria, and good for appropriateness of day of care. A patient day identified by ASSURE as potentially inappropriate would be twice as likely to be judged inappropriate by a reviewer as a randomly selected patient day. Review of the 10% of patient days identified as potentially inappropriate by ASSURE would identify approximately 21% of the inappropriate days of care. CONCLUSION: ASSURE is a clinically useful tool for screening adult acute care patients for inappropriate days of care, and promises to make a major contribution to reducing health care costs. The prognosis for successful routine clinical use is good.

Artificial Intelligence↗

A program for the user-independent computation of the correlation dimension and the largest Lyapunov exponent of heart rate dynamics from small data sets.

We propose a specially optimized computer program for the user-independent calculation of the correlation dimension D and the largest Lyapunov exponent L of heart rate dynamics on the basis of only 1024 electrocardiographically recorded RR intervals (heartbeat intervals). The validity of our program was established by analyzing a set of artificial standard signals. Our norm values of the correlation dimension (D = 5.37 +/- 0.62) and the largest Lyapunov exponent (L = 0.561 +/- 0.037 bits/beat) of RR dynamics, obtained from 79 healthy adults aged 26.3 +/- 4.8 years, were independent of gender and age; D and L correlated slightly with each other (Pearson correlation coefficient r = 0.26). Short-term reliability, tested for 25 of our subjects by two successive recordings, was fair: the intraclass correlation coefficients (ICCs) were 0.45 and 0.41 for D and L of RR dynamics, respectively. However, long-term reliability, tested for eight of our subjects by ten weekly recordings, was acceptable for L (ICC = 0.40) but not for D (ICC = 0.01). These results permit group comparisons on the basis of single measurements of L of RR dynamics. A reliable differentiation between young healthy individuals requires four measurements of L.

Adult↗

Use of an artificial neural network (ANN) for classifying nursing care needed, using incomplete input data.

BACKGROUND: In German nursing insurance, the act of classifying the client into four categories of disability is based on legally defined distinct criteria. When classifying deceased persons it is often impossible to collect all the required information. PRIMARY OBJECTIVE: We aimed to determine the ability of an artificial neural network (ANN) to calculate the category of disability, to investigate the response of the ANN to input items of different nature, quantity and data quality, and to estimate the minimum number of training data required. RESEARCH DESIGN: The investigation was conducted as a retrospective observational study. METHODS AND PROCEDURES: The analysis was based on routine records of 14000 adult clients of the nursing insurance. Several ANNs were trained, varying nature, number and quality of the input items as well as the size of the training data set. Each ANN's classification competence was tested on independent validation data, judging the ANN's conformance to the result of the individual expert assessment, using kappa statistics. MAIN RESULTS: Fed with all 30 input items available, the net classified 80% of cases correctly (weighted kappa = 0.78). Using three input items, weighted kappa was 0.63. Severe misclassification (deviation by more than one category in either direction) ranged between 0.2% (all 30 input items) and 3.7% (3/30 items). The less complete the individual input items were, the less accurate was the net's estimate. A 20% rate of missing values was well tolerated. A training set comprising 500 cases was adequate. CONCLUSIONS: The input item set inherits redundancy. The ANN's ability to correctly respond to subsets of input items makes it a powerful tool in quality control. In the categorization of deceased persons when only an incomplete input item set is available, the ANN can achieve satisfactory results.

Activities of Daily Living↗

Monitoring expert system performance using continuous user feedback.

OBJECTIVE: To evaluate the applicability of metrics collected during routine use to monitor the performance of a deployed expert system. METHODS: Two extensive formal evaluations of the GermWatcher (Washington University School of Medicine) expert system were performed approximately six months apart. Deficiencies noted during the first evaluation were corrected via a series of interim changes to the expert system rules, even though the expert system was in routine use. As part of their daily work routine, infection control nurses reviewed expert system output and changed the output results with which they disagreed. The rate of nurse disagreement with expert system output was used as an indirect or surrogate metric of expert system performance between formal evaluations. The results of the second evaluation were used to validate the disagreement rate as an indirect performance measure. Based on continued monitoring of user feedback, expert system changes incorporated after the second formal evaluation have resulted in additional improvements in performance. RESULTS: The rate of nurse disagreement with GermWatcher output decreased consistently after each change to the program. The second formal evaluation confirmed a marked improvement in the program's performance, justifying the use of the nurses' disagreement rate as an indirect performance metric. CONCLUSIONS: Metrics collected during the routine use of the GermWatcher expert system can be used to monitor the performance of the expert system. The impact of improvements to the program can be followed using continuous user feedback without requiring extensive formal evaluations after each modification. When possible, the design of an expert system should incorporate measures of system performance that can be collected and monitored during the routine use of the system.

Expert Systems↗

SuperStar: a knowledge-based approach for identifying interaction sites in proteins.

An empirical method for identifying interaction sites in proteins is described and validated. The method is based entirely on experimental information about non-bonded interactions occurring in small-molecule crystal structures. These data are used in the form of scatterplots that show the experimentally observed distribution of one functional group (the "contact group" or "probe") around another. A template molecule (e.g. a protein binding site) is broken down into structure fragments and the scatterplots, showing the distribution of a chosen probe around these structure fragments, are superimposed on the corresponding parts of the template. The scatterplots are then translated into a three-dimensional map that shows the propensity of the probe at different positions around the template molecule. The method is illustrated for l -arabinose-binding protein, complexed with l -arabinose and with d -fucose, and for dihydrofolate reductase complexed with methotrexate. The method is validated on 122 X-ray structures of protein-ligand complexes. For all the binding sites of these proteins, propensity maps are generated for four different probes: a charged NH+3nitrogen, a carbonyl oxygen, a hydroxyl oxygen and a methyl carbon atom. Next, the maps are compared with the experimentally observed positions of ligand atoms of these types. For 74% of these ligand atoms (84% of the solvent-inaccessible ones) the calculated propensity of the matching probe at the experimental positions is higher than expected by chance. For 68% of the atoms (82% of the solvent-inaccessible ones) the propensity of the matching probe is higher than that of the other three probes. These results indicate that the approach generally gives good predictions for protein-ligand interactions. The potential applications of the propensity maps range from an aid in manual docking and structure-based drug design to their use in pharmacophore development.

Artificial Intelligence↗

Mean regional cerebral blood flow images of normal subjects using technetium-99m-HMPAO by automated image registration.

UNLABELLED: The purpose of this study was twofold: to calculate relative uptake values for 99mTc-HMPAO in various regions of the normal brain after alignment and registration to a standard shape and size, and to validate the automated image registration (AIR) program for SPECT-to-SPECT transformation. METHODS: Thirty subjects took part in this study. Technetium-99m-HMPAO brain SPECT and x-ray-CT scans were acquired. SPECT images were normalized to an average activity of 100 counts/pixel. Intersubject accuracy was evaluated on brain images of 17 normal subjects (mean age = 64.9 +/- 8.7 yr). These images were aligned and registered to a standard size and shape with the help of AIR. Realigned images were overlaid on reference images to determine the overlap areas. Intrasubject accuracy was evaluated by realigning 20 degree rotated brain images with an index calculated as: overlap area/(overlap area + nonoverlap area). Anatomical variability between realigned target and reference images was evaluated by measurements on corresponding x-ray-CT scans, realigned using transformations that were established by the SPECT images. Realigned brain SPECT images of 30 normal subjects (mean age = 50.7 +/- 18.7 yr), including those subjects examined in the accuracy validation study, were used to generate mean and s.d. images. Images based on the mean value of each voxel (n = 30) were compared with other mean images prepared by the human brain atlas (HBA) standardization technique on a voxel-by-voxel basis to generate T maps. RESULTS: Accuracy indices were 0.98 +/- 0.006 and 0.99 +/- 0.002 for the intersubject and intrasubject evaluations, respectively. The maximum anatomical variability was 4.7 mm after realignment. Paired Student's t-test comparisons of mean HBA and AIR images revealed statistically significant differences for the deep white matter, pons and occipito-temporal regions. These differences could be explained by variation in the population being studied and the protocol for data handling by AIR and HBA. CONCLUSION: AIR aligns and registers brain SPECT images with acceptable accuracy, without the necessity of MRI or x-ray-CT scans.

Brain↗

A strategy for development of computerized critical care decision support systems.

It is not enough to merely manage medical information. It is difficult to justify the cost of hospital information systems (HIS) or intensive care unit (ICU) patient data management systems (PDMS) on this basis alone. The real benefit of an integrated HIS or PDMS is in decision support. Although there are a variety of HIS and ICU PDMS systems available there are few that provide ICU decision support. The HELP system at the LDS Hospital is an example of a HIS which provides decision support on many different levels. In the ICU there are decision support tools for antibiotic therapy, nutritional management, and management of mechanical ventilation. Computer protocols for the management of mechanical ventilation (respiratory evaluation, ventilation, oxygenation, weaning and extubation) in patients with adult respiratory distress syndrome ((ARDS) have already been developed and clinically validated at the LDS Hospital. These protocols utilize the bedside intensive care unit (ICU) computer terminal to prompt the clinical care team with therapeutic and diagnostic suggestions. The protocols (in paper flow diagram and computerized form) have been used for over 40,000 hours in more than 125 adult respiratory distress syndrome (ARDS) patients. The protocols controlled care for 94% of the time. The remainder of the time patient care was not protocol controlled was a result of the patient being in states not covered by current protocol logic (e.g. hemodynamic instability, or transport for X-Ray studies). 52 of these ARDS patients met extra corporal membrane oxygenation (ECMO) criteria. The survival of the ECMO criteria ARDS patients was 41%, four times that expected (9%) from historical data (p less than 0.0002).(ABSTRACT TRUNCATED AT 250 WORDS)

Attitude of Health Personnel↗

The development of a comprehensive, institution-based patient risk evaluation program: II. Validity and reliability of questionnaire data.

The accuracy of historical information derived from self-administered questionnaires must be confirmed. We report the results of studies conducted to assess the reliability and validity of data collected from a comprehensive cancer risk factor questionnaire developed at The University of Texas M.D. Anderson Cancer Center. A comparison of the basic demographic data of a randomly selected sample of 80 respondents and 70 nonrespondents revealed no fundamental ethnic or socioeconomic differences. We verified self-reported past illnesses, surgical procedures, and cancers by reviewing 72 patient charts, using stringent diagnostic criteria for verification. We noted substantial agreement between self-reported and documented illnesses and operations. With the exception of nine patients who misclassified metastatic disease, the verification of primary cancers was excellent. We determined reliability by interviewing 50 of these patients by telephone. Questions with a dichotomous outcome (e.g., smoking status) were reliably answered; however, those requiring quantification (e.g., amount of alcohol consumed) were less accurately reported on interview. While we recognize the limitations of self-administered questionnaires, we believe this program will develop into a comprehensive, standardized, easily accessible patient risk factor data base.

Cancer Care Facilities↗

Validity of computerized predictions of dentoskeletal and soft tissue profile changes after mandibular setback and maxillary impaction osteotomies.

The aim of the present investigation was to evaluate the validity of the predictions of a computerized cephalometric system (Dentofacial Planner) regarding dentoskeletal and soft tissue profile changes after mandibular setback and maxillary impaction osteotomies. Tracings of lateral cephalograms taken at the end of preoperative orthodontics (within 1 week before surgery) and approximately 1 year after the operation were digitized and entered into the Dentofacial Planner. For the mandibular setback group, the computerized predictions tended to place the mandible less posteriorly than the actual situation and to significantly underestimate the mandibular plane angle, the total anterior skeletal and soft tissue facial heights, the lower anterior skeletal facial height, and the upper lip height. In the maxillary impaction group, the prediction printouts significantly overestimated the total anterior soft tissue facial height, the upper lip height and the inclination and curvature of the lower lip and underestimated the soft tissue thickness in the regions of pogonion and point B.

Adolescent↗

Evaluation of Sleep Expert--a computer-aided decision support system for sleep disorders.

Sleep Expert--a medical decision support system--is a prototype program, with knowledge based on the International Classification of Sleep Disorders (1990). The goal of this project was to evaluate Sleep Expert. In the evaluation project the knowledge of the program was first validated. Three physicians, experts in sleep disorders, were asked to choose 10 typical patient cases with sleep disorders, and to write a description. They also made a diagnosis for each case. Next, each expert made a diagnosis of the cases supplied by the other experts. They were not given the original diagnosis. The 'right diagnosis' (so-called majority agreement) was determined from the three diagnoses. Then the diagnosis of each expert was compared with the 'right diagnosis'. Two physicians, not experts in sleep disorders, were asked to make a diagnosis by using Sleep Expert. Compared to the 'right diagnosis' the diagnoses of each user (non-expert physician) were correct to 63 and 70% of cases, which is quite a good result, although it does not reach the level of the expert physicians (> or = 87%). The functionality of Sleep Expert was studied by using a limited inquiry. On the basis of the user inquiry Sleep Expert provided a useful clinical tool for non-experts.

Diagnosis, Computer-Assisted↗

Risk-adjusting acute myocardial infarction mortality: are APR-DRGs the right tool?

OBJECTIVE: To determine if a widely used proprietary risk-adjustment system, APR-DRGs, misadjusts for severity of illness and misclassifies provider performance. DATA SOURCES: (1) Discharge abstracts for 116,174 noninstitutionalized adults with acute myocardial infarction (AMI) admitted to nonfederal California hospitals in 1991-1993; (2) inpatient medical records for a stratified probability sample of 974 patients with AMIs admitted to 30 California hospitals between July 31, 1990 and May 31, 1991. STUDY DESIGN: Using the 1991-1993 data set, we evaluated the predictive performance of APR-DRGs Version 12. Using the 1990/1991 validation sample, we assessed the effect of assigning APR-DRGs based on different sources of ICD-9-CM data. DATA COLLECTION/EXTRACTION METHODS: Trained, blinded coders reabstracted all ICD-9-CM diagnoses and procedures, and established the timing of each diagnosis. APR-DRG Risk of Mortality and Severity of Illness classes were assigned based on (1) all hospital-reported diagnoses, (2) all reabstracted diagnoses, and (3) reabstracted diagnoses present at admission. The outcome variables were 30-day mortality in the 1991-1993 data set and 30-day inpatient mortality in the 1990/1991 validation sample. PRINCIPAL FINDINGS: The APR-DRG Risk of Mortality class was a strong predictor of death (c = .831-.847), but was further enhanced by adding age and sex. Reabstracting diagnoses improved the apparent performance of APR-DRGs (c = .93 versus c = .87), while using only the diagnoses present at admission decreased apparent performance (c = .74). Reabstracting diagnoses had less effect on hospitals' expected mortality rates (r = .83-.85) than using diagnoses present at admission instead of all reabstracted diagnoses (r = .72-.77). There was fair agreement in classifying hospital performance based on these three sets of diagnostic data (K = 0.35-0.38). CONCLUSIONS: The APR-DRG Risk of Mortality system is a powerful risk-adjustment tool, largely because it includes all relevant diagnoses, regardless of timing. Although some late diagnoses may not be preventable, APR-DRGs appear suitable only if one assumes that none is preventable.

Adult↗

Linear models for the prediction of stature from foot and boot dimensions.

Estimation of stature from the dimensions of foot or shoeprints has considerable forensic value in developing descriptions of suspects from evidence at the crime scene and in corroborating height estimates from witnesses. This study extends the findings of previous researchers by exploring linear models with and without gender and race indicators, and by validating the most promising models on a large, recently collected military database. Boot size and outsole dimensions are also examined as predictors of stature. The results of this study indicate that models containing both foot length and foot breadth are significantly better than those containing only foot length. Models with race/gender indicators also perform significantly better than do models without race/gender indicators. However, the difference in performance is slight, and the availability of reliable gender and race information in most forensic situations is uncertain. Analogous results were obtained for models utilizing boot size/width and outsole length/width, and in this study these variables performed nearly as well as the foot dimensions themselves. Although the adjusted R2 values for these models clearly reflect a strong relationship between foot/boot length and stature, individual 95% prediction limits for even the best models are +/- 86 mm (3.4 in.). This suggests that models estimating stature from foot/shoe-prints may be useful in the development of subject descriptions early in a case but, because of their imprecision, may not always be helpful in excluding individual suspects from consideration.

Anthropometry↗

Validating a decision support system for anti-epileptic drug treatment. Part I: initiating anti-epileptic drug treatment.

In this contribution the validation of a prototype decision support system that implements a model of expertise for initiating anti-epileptic drug treatment is described. Since domain experts were of the opinion that prescribing was a rather straightforward process we used only one expert neurologist for knowledge elicitation. To determine the correctness of the system we intended to compare the contents of the system's prescriptions with the majority decision of three neurologists. Because of a large variation in prescribing a majority decision could not be obtained in many cases. Even a Delphi procedure did not yield a majority decision in a large number of cases. Therefore a consensus meeting was organised to discuss cases where discrepancies remained. In the process the participating neurologists formulated prescription guidelines. These guidelines were used as a reference to determine the correctness of the prescriptions of both the system and of the neurologists. The acceptability of all prescriptions for each case was rated by the two neurologists who did not write a prescription for that case. From both comparisons it could be concluded that the system was at least as good in prescribing as individual neurologists.

Adolescent↗

The Australian Incident Monitoring Study. Crisis management--validation of an algorithm by analysis of 2000 incident reports.

Anaesthetists are called upon to manage complex life-threatening crises at a moment's notice. As there is evidence that this may require cognitive tasking beyond the information-processing capacity of the human brain, it was decided to try and develop a generic crisis management algorithm analogous to the "Phase I" immediate response routine used by airline pilots. Such an algorithm, based on the mnemonic "COVER ABCD, A SWIFT CHECK", was developed and refined over 3 meetings, each attended by 60-100 anaesthetists and aviation psychologists. It was validated against 1301 relevant incidents among the first 2000 incidents reported to the Australian Incident Monitoring Study. It proved sufficiently robust and safe to recommend its general use as an initial response to any incident or crisis which occurs when a patient is breathing gas from an anesthetic machine. It requires a limited knowledge base and is easily learnt and rehearsed during the anaesthetist's working day. It will provide a functional diagnosis in over 99% of cases and will correct 62% of the problems in 40-60 seconds. In the remaining 37% it will allow the anaesthetist to proceed with a "sub-algorithm", confident in the knowledge that some important step has not been missed. In just over 30% of incidents this will be for a problem familiar to all anaesthetists (e.g. laryngospasm, bradycardia); in just over 6% it will be for a less common, more complex, but finite, set of problems (3% cardiac arrest, 1% air embolism, 1% anaphylaxis, 1% for the remaining desaturations); in less than 1% diagnosis and correction will require a more complex checklist (e.g. for malignant hyperthermia, pneumothorax). The next stage, the development of specific sub-algorithms and a structured team approach for ongoing problems, is in progress.

Algorithms↗