PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Molecular classification of human carcinomas by use of gene expression signatures.

Classification of human tumors according to their primary anatomical site of origin is fundamental for the optimal treatment of patients with cancer. Here we describe the use of large-scale RNA profiling and supervised machine learning algorithms to construct a first-generation molecular classification scheme for carcinomas of the prostate, breast, lung, ovary, colorectum, kidney, liver, pancreas, bladder/ureter, and gastroesophagus, which collectively account for approximately 70% of all cancer-related deaths in the United States. The classification scheme was based on identifying gene subsets whose expression typifies each cancer class, and we quantified the extent to which these genes are characteristic of a specific tumor type by accurately and confidently predicting the anatomical site of tumor origin for 90% of 175 carcinomas, including 9 of 12 metastatic lesions. The predictor gene subsets include those whose expression is typical of specific types of normal epithelial differentiation, as well as other genes whose expression is elevated in cancer. This study demonstrates the feasibility of predicting the tissue origin of a carcinoma in the context of multiple cancer classes.

Carcinoma↗

Automatic document classification of biological literature.

BACKGROUND: Document classification is a wide-spread problem with many applications, from organizing search engine snippets to spam filtering. We previously described Textpresso, a text-mining system for biological literature, which marks up full text according to a shallow ontology that includes terms of biological interest. This project investigates document classification in the context of biological literature, making use of the Textpresso markup of a corpus of Caenorhabditis elegans literature. RESULTS: We present a two-step text categorization algorithm to classify a corpus of C. elegans papers. Our classification method first uses a support vector machine-trained classifier, followed by a novel, phrase-based clustering algorithm. This clustering step autonomously creates cluster labels that are descriptive and understandable by humans. This clustering engine performed better on a standard test-set (Reuters 21578) compared to previously published results (F-value of 0.55 vs. 0.49), while producing cluster descriptions that appear more useful. A web interface allows researchers to quickly navigate through the hierarchy and look for documents that belong to a specific concept. CONCLUSION: We have demonstrated a simple method to classify biological documents that embodies an improvement over current methods. While the classification results are currently optimized for Caenorhabditis elegans papers by human-created rules, the classification engine can be adapted to different types of documents. We have demonstrated this by presenting a web interface that allows researchers to quickly navigate through the hierarchy and look for documents that belong to a specific concept.

Abstracting and Indexing↗

Application of fuzzy-classifier system to coronary artery disease and breast cancer.

This paper presents an application of a genetic-algorithm-based representation of fuzzy rules for the classification of coronary artery disease data and breast cancer data. The performance of this fuzzy classifier for classification of coronary artery disease and breast cancer data is evaluated. In this study the concept of fuzzy if-then has been applied of rules proposed by Ishibuchi et al. for a multi dimensional data classification problem which leads to higher classification power. The fitness value of each fuzzy if-then rule was determined by the numbers of correctly and wrongly classified training patterns for that rule. The classification power on real world data for coronary artery disease and breast cancer was thus demonstrated by computer simulations.

Algorithms↗

Lesion size quantification in SPECT using an artificial neural network classification approach.

An artificial neural network (ANN) has been developed to determine the size of lesions detected in single photon emission computed tomographic images. The network is the Learning Vector Quantizer and is trained to perform size quantification based on image neighborhoods extracted around the lesions. The ANN is compared to the optimal, Bayesian algorithm developed to perform the same task using the unreconstructed, projection data. The performance of the neural network is evaluated at two different noise levels. The Bayesian algorithm provides the upper bound for size quantification performance against which the ANN is compared. In the ideal case where the Bayesian algorithm has explicit knowledge of the underlying distributions, its performance is superior to that of the neural network. However, in the more realistic case where the distributions need to be estimated from the same learning sample the ANN was trained on, the two algorithms have comparable performances.

Algorithms↗

Developing a decision tree algorithm for the diagnosis of suspected spider bites.

OBJECTIVE: To develop a diagnostic algorithm (decision tree) to improve the ability to identify or predict medically important spider bites (funnel-web and redback spiders) from information about the circumstances and initial clinical effects of spider bites. METHODS: A dataset of definite spider bites with expert identification of all spiders was used from a previous Australia-wide prospective study. Spider bites were categorized as: big black spider (BBS), redback spider (RED) and other spider (OTH). Big black spider included funnel-web spiders (most medically significant), but also other spiders of similar appearance. Fifteen predictor variables were based on univariate analysis from previous studies and clinical experience. They included information about the circumstances and early clinical effects of bites. The data were analyzed using CART (Classification and Regression Trees), a 'decision tree' algorithm used to create a tree-like structure to describe a data set. RESULTS: Of 789 spider bites there were 49 (6.2%) bites by BBS, 68 (8.6%) bites by RED and 672 (85.2%) bites by OTH. A decision tree was developed that included six predictor variables (fang marks/bleeding; state/territory; local diaphoresis; month; time of day; and proximal or distal bite region). The decision tree accurately classified 47 out of the 49 (96%) BBS, and no funnel-web spiders were incorrectly classified (100% sensitivity). Two hundred and forty-four of 789 were classified as OTH and included no BBS. CONCLUSIONS: A decision tree based on a small amount of information about the circumstances and early clinical effects of spider bites safely predicted all funnel-web spider bites. Application of this algorithm would allow the early institution of appropriate treatment for funnel-web spider bites and the immediate discharge of 31% as other spider bites (reassurance only).

Algorithms↗

Neural network method to determine the vigilance levels of the central nervous system, related to occupational chronic chemical stress.

The effects of chronic toxic occupational factors and functional disorders of the central nervous system (CNS) in chemical industry were studied. These factors cause various stages of chronic chemical stress on the human CNS together with changes of the vigilance levels. On the basis of QEEG data analysis and psychometric tests we identified three stages of occupational chemical stress syndromes according to the CNS vigilance level (ordered from light to severe): hypersthenic syndrome, hyposthenic syndrome, and organic psychosyndrome. Each syndrome is characterized by specific changes in the QEEG data. A perceptron-based neural network was developed for the classification of the QEEG data to one of the above-mentioned syndrome classes. The data of 77 patients and 10 healthy subjects were selected to test the algorithm. Different combinations of the QEEG data as input features to the classifier were chosen. The most reliable classification was obtained when QEEG data measured during the visual stimulation of the CNS were used. However, sometimes the algorithm was unable to solve the classification problem, or it took a very long time to train the perceptron. In part, difficulties arose from using a perceptron-based algorithm, which can classify only linearly separable data.

Algorithms↗

Differential diagnosis of jaundice: a pocket diagnostic chart.

Based on extensive clinical and clinical chemical information (107 different items) from 1002 jaundiced patients, we developed a diagnostic algorithm which was evaluated on a test sample of another 110 jaundiced patients. A primary classification into categories of obstructive jaundice (probability of obstruction greater than or equal to 0.80), non-obstructive jaundice (probability of obstruction less than or equal to 0.20), and of doubtful causes of jaundice (probability of obstruction: 0.20-0.80) was attempted. Among 234 patients in the data base who were classified as obstructive, 220 (94%) proved to be so, as did 36 (97%) of 37 in the test sample. The corresponding figures for non-obstructive jaundice were 463 (96%) of 483 patients correctly classified in the data base and 47 (92%) of 51 patients in the test sample. Altogether 69% of the patients in the data base and 75% of those in the test sample were correctly classified, in 27% and 20% the cause of jaundice was doubtful, and only 4% and 5%, respectively, were misclassified. A slight majority of the patients in whom the algorithmic diagnoses were doubtful proved obstructive. A close correlation was found between the preliminary diagnoses made by the algorithm and by the clinicians. A secondary classification of the patients by the algorithm into benign versus malignant causes of obstructive jaundice performed equally well in the data base and the test sample.

Cholestasis↗

Relevance vector machine for optical diagnosis of cancer.

BACKGROUND AND OBJECTIVES: A probability-based, robust diagnostic algorithm is an essential requirement for successful clinical use of optical spectroscopy for cancer diagnosis. This study reports the use of the theory of relevance vector machine (RVM), a recent Bayesian machine-learning framework of statistical pattern recognition, for development of a fully probabilistic algorithm for autofluorescence diagnosis of early stage cancer of human oral cavity. It also presents a comparative evaluation of the diagnostic efficacy of the RVM algorithm with that based on support vector machine (SVM) that has recently received considerable attention for this purpose. STUDY DESIGN/MATERIALS AND METHODS: The diagnostic algorithms were developed using in vivo autofluorescence spectral data acquired from human oral cavity with a N(2) laser-based portable fluorimeter. The spectral data of both patients as well as normal volunteers, enrolled at Out Patient department of the Govt. Cancer Hospital, Indore for screening of oral cavity, were used for this purpose. The patients selected had no prior confirmed malignancy and were diagnosed of squamous cell carcinoma (SCC), Grade-I on the basis of histopathology of biopsy taken from abnormal site subsequent to acquisition of spectra. Autofluorescence spectra were recorded from a total of 171 tissue sites from 16 patients and 154 healthy squamous tissue sites from 13 normal volunteers. Of 171 tissues sites from patients, 83 were SCC and the rest were contralateral uninvolved squamous tissue. Each site was treated separately and classified via the diagnostic algorithm developed. Instead of the spectral data from uninvolved sites of patients, the data from normal volunteers were used as the normal database for the development of diagnostic algorithms. RESULTS: The diagnostic algorithms based on RVM were found to provide classification performance comparable to the state-of-the-art SVMs, while at the same time explicitly predicting the probability of class membership. The sensitivity and specificity towards cancer were up to 88% and 95% for the training set data based on leave- one-out cross validation and up to 91% and 96% for the validation set data. When implemented on the spectral data of the uninvolved oral cavity sites from the patients, it yielded a specificity of up to 91%. CONCLUSIONS: The Bayesian framework of RVM formulation makes it possible to predict the posterior probability of class membership in discriminating early SCC from the normal squamous tissue sites of the oral cavity in contrast to dichotomous classification provided by the non-Bayesian SVM. Such classification is very helpful in handling asymmetric misclassification costs like assigning different weights for having a false negative result for identifying cancer compared to false positive. The results further demonstrate that for comparable diagnostic performances, the RVM-based algorithms use significantly fewer kernel functions and do not need to estimate any hoc parameters associated with the learning or the optimization technique to be used. This implies a considerable saving in memory and computation in a practical implementation.

Algorithms↗

Potential of the genetic algorithm neural network in the assessment of gait patterns in ankle arthrodesis.

The aim of this study was to develop an empirical model of parameter-based gait data, based on an artificial neural network and a genetic algorithm, for the assessment of patients after ankle arthrodesis. Ground reaction force vectors were measured by force platforms during level walking. Nine force parameters expressed in percentage of body weight and their chronologic incidence of occurrence expressed in percentage of stance phase period were used in modeling. Ten healthy persons and ten patients who had solid arthrodesis of the ankle were recruited in this study for developing the model. By applying the genetic algorithm neural network, the percentage of correct classification was 98.8% and the subset of discriminant parameters was be reduced to 9 out of 18. These key parameters were mainly related to the loading response and propulsive phase. This indicates that there was a reduction in the abilities in cushion impact and push off in the patients after ankle arthrodesis. Finally, the relative distance (Dr) was defined in this study and used in two new patients' examinations to demonstrate its clinical utility.

Adult↗

Image analysis and pattern recognition for computer supported skin tumor diagnosis.

A new approach to computer supported recognition of melanoma and naevocytic naevi based on high resolution skin surface profiles is presented. Profiles are generated by sampling an area of 4 x 4 mm2 at a resolution of 125 sample points per mm with a laser profilometer at a vertical resolution of 0.1 micron. With image analysis algorithms Haralick's texture parameters, Fourier features and features based on fractal analysis are extracted. Genetic algorithms are employed successfully to select good feature subsets for the following classification process. As quality measure for feature subsets, the error rate of the nearest neighbor classifier estimated with the leaving-one-out method is used. Classification is performed with feed forward back-propagation network and the nearest neighbor classifier. Classification performance of the neural classifier is optimized using different topologies, learning parameters and pruning algorithms. The best neural classifier achieved an error rate of 4.5% and was found after network pruning. The best result with an error rate of 2.3% was obtained with the nearest neighbor classifier.

Algorithms↗

Markers of adenocarcinoma characteristic of the site of origin: development of a diagnostic algorithm.

PURPOSE: Patients with metastatic adenocarcinoma of unknown origin are a common clinical problem. Knowledge of the primary site is important for their management, but histologically, such tumors appear similar. Better diagnostic markers are needed to enable the assignment of metastases to likely sites of origin on pathologic samples. EXPERIMENTAL DESIGN: Expression profiling of 27 candidate markers was done using tissue microarrays and immunohistochemistry. In the first (training) round, we studied 352 primary adenocarcinomas, from seven main sites (breast, colon, lung, ovary, pancreas, prostate and stomach) and their differential diagnoses. Data were analyzed in Microsoft Access and the Rosetta system, and used to develop a classification scheme. In the second (validation) round, we studied 100 primary adenocarcinomas and 30 paired metastases. RESULTS: In the first round, we generated expression profiles for all 27 candidate markers in each of the seven main primary sites. Data analysis led to a simplified diagnostic panel and decision tree containing 10 markers only: CA125, CDX2, cytokeratins 7 and 20, estrogen receptor, gross cystic disease fluid protein 15, lysozyme, mesothelin, prostate-specific antigen, and thyroid transcription factor 1. Applying the panel and tree to the original data provided correct classification in 88%. The 10 markers and diagnostic algorithm were then tested in a second, independent, set of primary and metastatic tumors and again 88% were correctly classified. CONCLUSIONS: This classification scheme should enable better prediction on biopsy material of the primary site in patients with metastatic adenocarcinoma of unknown origin, leading to improved management and therapy.

Adenocarcinoma↗

Multiclass Decision Forest--a novel pattern recognition method for multiclass classification in microarray data analysis.

The wealth of knowledge imbedded in gene expression data from DNA microarrays portends rapid advances in both research and clinic. Turning the prodigious and noisy data into knowledge is a challenge to the field of bioinformatics, and development of classifiers using supervised learning techniques is the primary methodological approach for clinical application using gene expression data. In this paper, we present a novel classification method, multiclass Decision Forest (DF), that is the direct extension of the two-class DF previously developed in our lab. Central to DF is the synergistic combining of multiple heterogenic but comparable decision trees to reach a more accurate and robust classification model. The computationally inexpensive multiclass DF algorithm integrates gene selection and model development, and thus eliminates the bias of gene preselection in crossvalidation. Importantly, the method provides several statistical means for assessment of prediction accuracy, prediction confidence, and diagnostic capability. We demonstrate the method by application to gene expression data for 83 small round blue-cell tumors (SRBCTs) samples belonging to one of four different classes. Based on 500 runs of 10-fold crossvalidation, tumor prediction accuracy was approximately 97%, sensitivity was approximately 95%, diagnostic sensitivity was approximately 91%, and diagnostic accuracy was approximately 99.5%. Among 25 genes selected to distinguish tumor class, 12 have functional information in the literature implicating their involvement in cancer. The four types of SRBCTs samples are also distinguishable in a clustering analysis based on the expression profiles of these 25 genes. The results demonstrated that the multiclass DF is an effective classification method for analysis of gene expression data for the purpose of molecular diagnostics.

Carcinoma, Small Cell↗

Diagnostic classification of urothelial cells in urine cytology specimens using exclusively spectral information.

BACKGROUND: Although cytologic evaluation of urine specimens is a standard procedure in the diagnosis and follow-up of bladder carcinoma, its sensitivity and specificity are low. Cytopathologic diagnoses are driven primarily by spatial relations or morphology. Although color enhances the pathologist's perception of the specimen, spectral information plays a minimal role in diagnostic processes. Recently, methods have been developed to capture and analyze spectral information from clinical specimens. In the current study, the authors determined the classification value of spectral information by testing its ability to discriminate between malignant and benign urothelial cells in cytology specimens. METHODS: Multiple images of benign urothelial cells (n = 39) and urothelial carcinoma cells (n = 35) were collected at serial wavelengths using a liquid crystal tunable optical filter and composited into a mosaic using ENVI (Environment for Visualizing Images) software. Through minimum noise fractionation and principal component analysis, the spectral information in the mosaic was compressed into a 29-dimensional scatter plot. The data generated were analyzed using visual and spectral end member extraction on both the original data set and a second independent data set (test set). RESULTS: One area of spectral clustering in the scatter plot segmented with carcinoma cells exclusively (100% specific), but was not present in every cell (approximately 50%), which may indicate that these spectral profiles are present in a subpopulation of malignant cells or at specific points of their cell cycle. Using ENVI algorithms, the authors found that a particular classification spectrum (end member 9) and its closest relatives identified malignant cell clusters, with a sensitivity and specificity that reached 82% and 81%, respectively. To validate this mechanism in a test set, a second mosaic comprised of 15 benign and 15 malignant clusters was analyzed using end member 9, resulting in a combined sensitivity and specificity of 73%. CONCLUSIONS: The results of the current study demonstrate that spectral information, in the complete absence of morphologic or spatial information, allows discrimination of benign and malignant urothelial cells in routine urine cytology specimens. The authors believe that this novel technology, combined with spatial analysis, has the potential to serve as an ancillary test for improved detection of bladder carcinoma.

Cytodiagnosis↗

A PC based neural network algorithm for measurement of heart rate variability.

Heart Rate Variability has recently been shown as a viable index to predict sudden cardiac death. The goal of this research is to investigate the use of neural network technique to classify detected QRS complexes into normal and abnormal ones. A single layer perceptron neural network is used for this QRS pattern learning and classification. Results with real data showed that the algorithm gives a 99% correct QRS detection rate.

Algorithms↗

Algorithms for the diagnosis and management of idiopathic anaphylaxis.

A major approach to the prevention of anaphylaxis due to exposure to an allergen is avoidance of the allergens. In cases of idiopathic anaphylaxis (IA), the method of management must first be pharmacologic control of the acute episode of IA and then the prevention of recurrent episodes. The appropriate management requires diagnosis, classification, and immediate initiation of a treatment regimen that will control and then induce a remission in patients with IA of the frequent (F) type. The classification and management of IA are reviewed and algorithms are presented for initial and long-term management of patients with IA.

Algorithms↗

New objective classification system for nuclear opacification.

We have developed an autonomous objective classification scheme for degree of nuclear opacification. The algorithm was developed by using a series of color 35-mm slides acquired with a Topcon photo slit-lamp microscope and use of standard camera settings. The photographs were digitized, and first, and second-order gray-level statistics were extracted from within circular regions of the nucleus. Classifications of severity were performed by using these features as input to a neural network. Training versus classification performance was tested by using photographs of different eyes, and test/retest classification reproducibility was evaluated by using paired photographs of the same eyes. We demonstrate good performance of the classifier against subjective assessments rendered by the Wilmer grading system [Invest. Ophthalmol. Visual Sci. 29, 73 (1988)] and markedly better test/retest reproducibility.

Algorithms↗

[Numerical taxonomy of the genus Desulfovibrio by group analysis].

The Desulfovibrio genus has a particular interest because it includes the microorganisms connected with the corrosion produced microbiologically. The taxonomy of the genus shows disadvantages due to its metabolical and physiological characteristics. In this paper, 14 strains of the Desulfovibrio type were studied from the metabolical point of view. Numeric taxonomy was carried out according to the Group Analysis method, using and comparing the change possibilities of the method. The Consensus Method was also applied. The results obtained indicate a low metabolic activity of the strains with regard to the number of compounds which can be used as energy source. The taxonomic method showed a better structure with more clear divisions, corresponding to Simple Matching coefficient (which coincides with other symmetric coefficients and with the distance coefficient) with average bond (UPGMA). It is estimated that the present classification will vary in time with new strains with different metabolic characteristics. The two groups of bacteria correspond to those with more and less degrading ability.

Algorithms↗

Accuracy-based learning classifier systems: models, analysis and applications to classification tasks.

Recently, Learning Classifier Systems (LCS) and particularly XCS have arisen as promising methods for classification tasks and data mining. This paper investigates two models of accuracy-based learning classifier systems on different types of classification problems. Departing from XCS, we analyze the evolution of a complete action map as a knowledge representation. We propose an alternative, UCS, which evolves a best action map more efficiently. We also investigate how the fitness pressure guides the search towards accurate classifiers. While XCS bases fitness on a reinforcement learning scheme, UCS defines fitness from a supervised learning scheme. We find significant differences in how the fitness pressure leads towards accuracy, and suggest the use of a supervised approach specially for multi-class problems and problems with unbalanced classes. We also investigate the complexity factors which arise in each type of accuracy-based LCS. We provide a model on the learning complexity of LCS which is based on the representative examples given to the system. The results and observations are also extended to a set of real world classification problems, where accuracy-based LCS are shown to perform competitively with respect to other learning algorithms. The work presents an extended analysis of accuracy-based LCS, gives insight into the understanding of the LCS dynamics, and suggests open issues for further improvement of LCS on classification tasks.

Algorithms↗