PubMed Health⌕ Search

Biomedical subjects

L Ohno-Machado

Publications and source records attributed to L Ohno-Machado.

At least 19 recordsLinked to original sources

A genetic algorithm approach to multi-disorder diagnosis.

One of the common limitations of expert systems for medical diagnosis is that they make an implicit assumption that multiple disorders do not co-occur in a single patient. The need for this simplifying assumption stems from the fact that finding minimal sets of disorders that cover all symptoms for a given patient is generally computationally intractable (NP-hard). In this paper, we explain the need for performing multi-disorder diagnosis, review previous approaches, formulate the problem using set theory notation, and propose the use of a search method based on a genetic algorithm. We test the algorithm and compare it to another approach using a simple example. The genetic algorithm performs well independently of the order of symptoms, and has the potential to perform multi-disorder diagnosis using existing or newly developed knowledge bases.

Algorithms↗

Using Boolean reasoning to anonymize databases.

This paper investigates how Boolean reasoning can be used to make the records in a database anonymous. In a medical setting, this is of particular interest due to privacy issues and to prevent the possible misuse of confidential information. As electronic medical records and medical data repositories get more common and widespread, the issue of making sensitive data anonymous becomes increasingly important. A theoretically well-founded algorithm is proposed that via cell suppression can be used to make a database anonymous before releasing or sharing it to the outside world. The degree of anonymity can be tailored according to the specific needs of the recipient, and according to the amount of trust we place in the recipient. Furthermore, the required measure of anonymity can be specified as far down as to the individual objects in the database. The algorithm can also be used for anonymization relative to a particular piece of information, effectively blocking deterministic inferences about sensitive database fields.

Adult↗

Evaluating variable selection methods for diagnosis of myocardial infarction.

This paper evaluates the variable selection performed by several machine-learning techniques on a myocardial infarction data set. The focus of this work is to determine which of 43 input variables are considered relevant for prediction of myocardial infarction. The algorithms investigated were logistic regression (with stepwise, forward, and backward selection), backpropagation for multilayer perceptrons (input relevance determination), Bayesian neural networks (automatic relevance determination), and rough sets. An independent method (self-organizing maps) was then used to evaluate and visualize the different subsets of predictor variables. Results show good agreement on some predictors, but also variability among different methods; only one variable was selected by all models.

Algorithms↗

A framework and tools for authoring, editing, documenting, sharing, searching, navigating, and executing computer-based clinical guidelines.

With the spread of managed care and integrated delivery networks, an increased emphasis has been placed on the cost-effectiveness of clinical practices. The need has been recognized to use guidelines to support education, and to integrate them into clinical practice. A specification for guideline representation that would facilitate computer-based clinical guideline sharing has been developed by the InterMed Collaboratory. Called GLIF (GuideLine Interchange Format), this specification and its proposed extensions have been the basis for our implementation of a framework and suite of integrated software tools for guideline authoring and editing, packaging in XML, Internet distribution, navigation, eligibility determination, and automatic execution.

Eligibility Determination↗

Decision support for clinical trial eligibility determination in breast cancer.

We have developed a system for clinical trial eligibility determination where patients or primary care providers can enter clinical information about a patient and obtain a ranked list of clinical trials for which the patient is likely to be eligible. We used clinical trial eligibility information from the National Cancer Institute's Physician Data Query (PDQ) database. We translated each free-text eligibility criterion into a machine executable statement using a derivation of the Arden Syntax. Clinical trial protocols were then structured as collections of these eligibility criteria using XML. The application compares the entered patient information against each of the eligibility criteria and returns a numerical score. Results are displayed in order of likelihood of match. We have tested our system using all phase II and III clinical trials for treatment of metastatic breast cancer found in the PDQ database. Preliminary results are encouraging.

Algorithms↗

A genetic algorithm to select variables in logistic regression: example in the domain of myocardial infarction.

Actual use of regression models in clinical practice depends on model simplicity. Reducing the number of variables in a model contributes to this goal. The quality of a particular selection of variables for a logistic regression model can be defined in terms of the number of variables selected and the model's discriminatory performance, as measured by the area under the ROC curve. A genetic algorithm was applied to search for the best variable combinations for modeling presence of myocardial infarction in a data set of patients with chest pain. Using an external validation set, the resulting model was compared with models constructed with standard backward, forward and stepwise methods of variable selection. The improvement in discriminatory ability yielded by the genetic algorithm variable selection method was statistically significant (p < 0.02).

Algorithms↗

Diagnosing breast cancer from FNAs: variable relevance in neural network and logistic regression models.

We compared the selection of variables for building a classification model for the diagnosis of breast cancer using neural networks and logistic regression. A set of 460 cases was used to build neural network and logistic regression models that classify cell samples obtained by fine-needle aspiration (FNA) as malignant or benign, depending on nine pathology features. Variables selected by a step down logistic regression model were compared to those selected by a measure of relevance derived from neural network weights. Since both types of models resulted in similar predictive accuracy, we expected approximately the same variables to be selected. The variables with the highest relevance values for the neural network models corresponded to those of high significance in univariate logistic regression models, but were not the ones selected in the step down procedure of multivariate models. Variable relevance based on weights for neural network models does not seem to be a consistent index of the importance of that variable for multivariate models such as logistic regression.

Analysis of Variance↗

Improving machine learning performance by removing redundant cases in medical data sets.

Neural network models and other machine learning methods have successfully been applied to several medical classification problems. These models can be periodically refined and retrained as new cases become available. Since training neural networks by backpropagation is time consuming, it is desirable that a minimum number of representative cases be kept in the training set (i.e., redundant cases should be removed). The removal of redundant cases should be carefully monitored so that classification performance is not significantly affected. We made experiments on data removal on a data set of 700 patients suspected of having myocardial infarction and show that there is no statistical difference in classification performance (measured by the differences in areas under the ROC curve on two previously unknown sets of 553 and 500 cases) when as many as 86% of the cases are randomly removed. A proportional reduction in the amount of time required to train the neural network model is achieved.

Area Under Curve↗

Comparison of multiple prediction models for ambulation following spinal cord injury.

Few studies have properly compared predictive performance of different models using the same medical data set. We developed and compared 3 models (logistic regression, neural networks, and rough sets) in the in prediction of ambulation at hospital discharge following spinal cord injury. We used the multi-center Spinal Cord Injury Model System database. All models performed well and had areas under the receiver operating characteristic curve in the 0.88-0.91 range. All models had sensitivity, specificity, and accuracy greater than 80% at ideal thresholds. The performance of neural network and logistic regression methods was not statistically different (p = 0.48). The rough sets classifier performed statistically worse than either the neural network or logistic regression models (p-values 0.002 and 0.015 respectively).

Acute Disease↗

Building manageable rough set classifiers.

An interesting aspect of techniques for data mining and knowledge discovery is their potential for generating hypotheses by discovering underlying relationships buried in the data. However, the set of possible hypotheses is often very large and the extracted models may become prohibitively complex. It is therefore typically desirable to only consider the "strongest" hypotheses, so that smaller models can be obtained that also retain good classificatory capabilities. This paper outlines how rule-based classifiers based on rough set theory and Boolean reasoning that are both small and perform well can be developed. Applied to a real-world medical dataset, the final models are shown to exhibit good performance using only a subset of the available information. Furthermore, the number of resulting rules is low and enables practical a posteriori inspection and interpretation of the models.

Classification↗

A comparison of Cox proportional hazards and artificial neural network models for medical prognosis.

Modeling survival of populations and establishing prognoses for individual patients are important activities in the practice of medicine. For patients with diseases that may extend for several years, in particular, accurate assessment of survival probabilities is essential. New methods, such as neural networks, have been used increasingly to model disease progression. Their advantages and disadvantages, when compared to statistical methods such as Cox proportional hazards, have seldom been explored in real-world data. In this study, we compare the performances of a Cox model and a neural network model that are used as prognostic tools for a set of people living with AIDS. We modeled disease progressions for patients who had AIDS (according to the 1993 CDC definition) in a set of 588 patients in California, using data from the ATHOS project. We divided the study population into 10 training and 10 test sets and evaluated the prognostic accuracy of a Cox proportional hazards model and of a neural network model by determining sensitivities, specificities, positive and negative predictive values for an arbitrary threshold (0.5), and the areas under the receiver operating characteristics (ROC) curves that utilized all possible thresholds for intervals of 1 yr following the diagnosis of AIDS. There was no evidence that the Cox model performed better than did the neural network model or vice versa, but the former method had the advantage of providing some insight on which variables were most influential for prognosis. Nevertheless, it is likely that the assumptions required by the Cox model may not be satisfied in all data sets, justifying the use of neural networks in certain cases.

Acquired Immunodeficiency Syndrome↗

Sequential versus standard neural networks for pattern recognition: an example using the domain of coronary heart disease.

The goal of this study was to compare standard and sequential neural network models for recognition of patterns of disease progression. Medical researchers who perform prognostic modeling usually oversimplify the problem by choosing a single point in time to predict outcomes (e.g. death in 5 years). This approach not only fails to differentiate patterns of disease progression, but also wastes important information that is usually available in time-oriented research data bases. The adequate use of sequential neural networks can improve the performance of prognostic systems if the interdependencies among prognoses at different intervals of time are explicitly modeled. In such models, predictions for a certain interval of time (e.g. death within 1 year) are influenced by predictions made for other intervals, and prognostic survival curves that provide consistent estimates for several points in time can be produced. We developed a system of neural network models that makes use of time-oriented data to predict development of coronary heart disease (CHD), using a set of 2594 patients. The output of the neural network system was a prognostic curve representing survival without CHD, and the inputs were the values of demographic, clinical, and laboratory variables. The system of neural networks was trained by backpropagation and its results were evaluated in test sets of previously unseen cases. We showed that, by explicitly modeling time in the neural network architecture, the performance of the prognostic index, measured by the area under the receiver operating characteristic (ROC) curve, was significantly improved (p < 0.05).

Adult↗

A virtual repository approach to clinical and utilization studies: application in mammography as alternative to a national database.

A national mammography database was proposed, based on a centralized architecture for collecting, monitoring, and auditing mammography data. We have developed an alternative architecture relying on Internet-based distributed queries to heterogeneous databases. This architecture creates a "virtual repository", or a federated database which is constructed dynamically, for each query and makes use of data available in legacy systems. It allows the construction of custom-tailored databases at individual sites that can serve the dual purposes of providing data (a) to researchers through a common mammography repository and (b) to clinicians and administrators at participating institutions. We implemented this architecture in a prototype system at the Brigham and Women's Hospital to show its feasibility. Common queries are translated dynamically into database-specific queries, and the results are aggregated for immediate display or download by the user. Data reside in two different databases and consist of structured mammography reports, coded per BIRADS Standardized Mammography Lexicon, as well as pathology results. We prospectively collected data on 213 patients, and showed that our system can perform distributed queries effectively. We also implemented graphical exploratory analysis tools to allow visualization of results. Our findings indicate that the architecture is not only feasible, but also flexible and scaleable, constituting a good alternative to a national mammography database.

Computer Communication Networks↗

Sequential use of neural networks for survival prediction in AIDS.

Prognostic assessment of patients is a key part of medical care. Although neural networks can be used to model survival, their accuracy has been limited for a variety of factors, including (1) the lack of data balance in certain intervals and (2) the lack of representation of temporal dependencies in the network architecture. Both problems can be solved with the use of sequential neural networks, which establish predictions for a certain time point and then use these predictions to produce survival estimates for other time points. If the sequence of models is adequate, sequential neural networks produce more accurate estimates of survival than standard neural networks, as shown in this example in the domain of AIDS. Assessments of survival in one, two, three, five and six years become more accurate (as measured by the areas under the ROC curves) when initial predictions of survival in four years are used in a sequential neural network model.

Acquired Immunodeficiency Syndrome↗

A comparison of two computer-based prognostic systems for AIDS.

We compare the performances of a Cox model and a neural network model that are used as prognostic tools for a cohort of people living with AIDS. We modeled disease progression for patients who had AIDS (according to the 1993 CDC definition) in a cohort of 588 patients in California, using data from the ATHOS project. We divided the study population into 10 training and 10 test sets and evaluated the prognostic accuracy of a Cox proportional hazards model and of a neural network model by determining the number of predicted deaths, the sensitivities, specificities, positive predictive values, and negative predictive values for intervals of one year following the diagnosis of AIDS. For the Cox model, we further tested the agreement between a series of binary observations, representing death in one, two, and three years, and a set of estimates which define the probability of survival for those intervals. Both models were able to provide accurate numbers on how many patients were likely to die at each interval, and reasonable individualized estimates for the two- and three-year survival of a given patient, but failed to provide reliable predictions for the first year after diagnosis. There was no evidence that the Cox model performed better than did the neural network model or vice-versa, but the former method had the advantage of providing some insight on which variables were most influential for prognosis. Nevertheless, it is likely that the assumptions required by the Cox model may not be satisfied in all data sets, justifying the use of neural networks in certain cases.

Acquired Immunodeficiency Syndrome↗

Hierarchical neural networks for survival analysis.

Neural networks offer the potential of providing more accurate predictions of survival time than do traditional methods. Their use in medical applications has, however, been limited, especially when some data is censored or the frequency of events is low. To reduce the effect of these problems, we have developed a hierarchical architecture of neural networks that predicts survival in a stepwise manner. Predictions are made for the first time interval, then for the second, and so on. The system produces a survival estimate for patients at each interval, given relevant covariates, and is able to handle continuous and discrete variables, as well as censored data. We compared the hierarchical system of neural networks with a nonhierarchical system for a data set of 428 AIDS patients. The hierarchical model predicted survival more accurately than did the nonhierarchical (although both had low sensitivity). The hierarchical model could also learn the same patterns in less than half the time required by the nonhierarchical model. These results suggest that the use of hierarchical systems is advantageous when censored data is present, the number of events is small, and time-dependent variables are necessary.

Acquired Immunodeficiency Syndrome↗

Identification of low frequency patterns in backpropagation neural networks.

Although neural networks have been widely applied to medical problems in recent years, their applicability has been limited for a variety of reasons. One of these barriers has been the inability to discriminate rare classes of solutions (i.e., the identification of categories that are infrequent). In this article, I demonstrate that a system of hierarchical neural networks (HNN) can overcome the problem of recognizing low frequency patterns, and therefore can improve the prediction power of neural-network systems. HNN are designed according to a divide-and-conquer approach: Triage networks are able to discriminate supersets that contain the infrequent pattern, and these supersets are then used by Specialized networks, which discriminate the infrequent pattern from the other ones in the superset. The supersets that are discriminated by the Triage networks are based on pattern similarity. The application of multilayered neural networks in more than one step allows the prior probability of a given pattern to increase at each step, provided that the predictive power of the network at the previous level is high. The method has been applied to one artificial set and one real set of data. In the artificial set, the distribution of the patterns was known and no noise was present. In this experiment, the HNN provided better discrimination than a standard neural network for all classes. In a real data set of nine thousand patients who were suspected of having thyroid disorders, the HNN also provided higher sensitivity than its corresponding standard neural network (without a corresponding decay in specificity) given the same time constraints.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

Prognostic classification for AIDS patients in Brazil.

We studied the survival rates and the prognostic variables that corresponded to death during hospitalization in 312 AIDS CDC group IV patients in São Paulo. Discriminant analysis proved to be a good tool to perform the exploratory data analysis that guided the survival analysis groups. It selected nine variables that were important in the progress of the disease: age, time elapsed from the first manifestations of the disease, gender, infection by helminths, number of risk groups to which the patient belonged, number of infections by fungi, history of transfusion, presence of esophageal candidiasis, and infection by Cryptosporidium sp. Although some of these variables may be of limited importance in developed countries, and some variables that we expected to be important were not present in the final discriminant function, we believe these results may guide future research in prognosis of death in hospitalized AIDS patients.

AIDS-Related Opportunistic Infections↗