PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

PCAS--a precomputed proteome annotation database resource.

BACKGROUND: Many model proteomes or "complete" sets of proteins of given organisms are now publicly available. Much effort has been invested in computational annotation of those "draft" proteomes. Motif or domain based algorithms play a pivotal role in functional classification of proteins. Employing most available computational algorithms, mainly motif or domain recognition algorithms, we set up to develop an online proteome annotation system with integrated proteome annotation data to complement existing resources. RESULTS: We report here the development of PCAS (ProteinCentric Annotation System) as an online resource of pre-computed proteome annotation data. We applied most available motif or domain databases and their analysis methods, including hmmpfam search of HMMs in Pfam, SMART and TIGRFAM, RPS-PSIBLAST search of PSSMs in CDD, pfscan of PROSITE patterns and profiles, as well as PSI-BLAST search of SUPERFAMILY PSSMs. In addition, signal peptide and TM are predicted using SignalP and TMHMM respectively. We mapped SUPERFAMILY and COGs to InterPro, so the motif or domain databases are integrated through InterPro. PCAS displays table summaries of pre-computed data and a graphical presentation of motifs or domains relative to the protein. As of now, PCAS contains human IPI, mouse IPI, and rat IPI, A. thaliana, C. elegans, D. melanogaster, S. cerevisiae, and S. pombe proteome.PCAS is available at http://pak.cbi.pku.edu.cn/proteome/gca.php CONCLUSION: PCAS gives better annotation coverage for model proteomes by employing a wider collection of available algorithms. Besides presenting the most confident annotation data, PCAS also allows customized query so users can inspect statistically less significant boundary information as well. Therefore, besides providing general annotation information, PCAS could be used as a discovery platform. We plan to update PCAS twice a year. We will upgrade PCAS when new proteome annotation algorithms identified.

Algorithms↗

A novel hybrid GA/RBFNN technique for protein sequences classification.

A novel hybrid genetic algorithm (GA)/radial basis function neural network (RBFNN) technique, which selects features from the protein sequences and trains the RBF neural network simultaneously, is proposed in this paper. Experimental results show that the proposed hybrid GA/RBFNN system outperforms the BLAST and the HMMer.

Algorithms↗

Bidirectional classification procedures: double tree and double cluster.

Partition based prediction rules as derived in CART, RECPAM or CBR, relating multivariate predictor variables X to a response variable Y, may be extended to bidirectional classification procedures. Tree or clustering algorithms are simultaneously applied to X and Y resulting in a double partition classification table that should be able to detect existing characteristic patterns of predictor and response values and relate them to each other.

Algorithms↗

Pattern recognition in gene expression profiling using DNA array: a comparative study of different statistical methods applied to cancer classification.

Large-scale parallel measurements of the expression of many thousands genes are now available with high-density array made with collections of cDNA fragments, or oligonucleotide corresponding to different transcripts. These technologies have been applied to cancer investigations since the availability of such a large number of markers makes DNA array a powerful diagnostic tool for tumour and patient classification. Over the last two years, a series of computational tools have been developed for the analysis of different aspects of gene profiling. Our work tries to compare a series of supervised statistical techniques on the basis of their ability to correctly classify different types of tumours. A simulation approach was initially used to control the huge source of variation among and between patients, and to evaluate the ability of algorithms to classify tumours in relation to different types of experimental variables. Different techniques for reduction of data dimension were then added to the discriminant analysis and compared according to their ability to capture the main genetic information. The simulation results have been tested by applying the selected classification algorithms to two experimental microarray datasets of human cancers, and by measuring the correspondent rates of misclassification. Our analyses identify in these datasets a series of genes principally involved in tumour characterization. The functional role of these discriminant transcripts is discussed.

Algorithms↗

Risk- and response-based classification of childhood B-precursor acute lymphoblastic leukemia: a combined analysis of prognostic markers from the Pediatric Oncology Group (POG) and Children's Cancer Group (CCG).

The Children's Cancer Group (CCG) and the Pediatric Oncology Group (POG) joined to form the Children's Oncology Group (COG) in 2000. This merger allowed analysis of clinical, biologic, and early response data predictive of event-free survival (EFS) in acute lymphoblastic leukemia (ALL) to develop a new classification system and treatment algorithm. From 11 779 children (age, 1 to 21.99 years) with newly diagnosed B-precursor ALL consecutively enrolled by the CCG (December 1988 to August 1995, n=4986) and POG (January 1986 to November 1999, n=6793), we retrospectively analyzed 6238 patients (CCG, 1182; POG, 5056) with informative cytogenetic data. Four risk groups were defined as very high risk (VHR; 5-year EFS, 45% or below), lower risk (5-year EFS, at least 85%), and standard and high risk (those remaining in the respective National Cancer Institute [NCI] risk groups). VHR criteria included extreme hypodiploidy (fewer than 44 chromosomes), t(9;22) and/or BCR/ABL, and induction failure. Lower-risk patients were NCI standard risk with either t(12;21) (TEL/AML1) or simultaneous trisomies of chromosomes 4, 10, and 17. Even with treatment differences, there was high concordance between the CCG and POG analyses. The COG risk classification scheme is being used for division of B-precursor ALL into lower- (27%), standard- (32%), high- (37%), and very-high- (4%) risk groups based on age, white blood cell (WBC) count, cytogenetics, day-14 marrow response, and end induction minimal residual disease (MRD) by flow cytometry in COG trials.

Adolescent↗

Feature selection for descriptor based classification models. 2. Human intestinal absorption (HIA).

We show that the topological polar surface area (TPSA) descriptor and the radial distribution function (RDF) applied to electronic and steric atom properties, like the conjugated electrotopological state (CETS), are the most relevant features/descriptors for predicting the human intestinal absorption (HIA) out of a large set of 2934 features/descriptors. A HIA data set with 196 molecules with measured HIA values and 2934 features/descriptors were calculated using JOELib and MOE. We used an adaptive boosting algorithm to solve the binary classification problem (AdaBoost.M1) and Genetic Algorithms based on Shannon Entropy Cliques (GA-SEC) variants as hybrid feature selection algorithms. The selection of relevant features was applied with respect to the generalization ability of the classification model, avoiding a high variance for unseen molecules (overfitting).

Humans↗

A robust algorithm for ratio estimation in two-color microarray experiments.

The reliability of the algorithms for ratio estimation in two-color microarray image analysis is very important, as these ratios build up the primary source of information for the subsequent analytical procedures (normalization, clustering, classification, etc). Although various algorithms already exist, there is still a need to develop procedures having higher levels of accuracy and robustness. We present a statistical procedure for the detection and removal of aberrant pixels in two-color microarray images. It is based on a linear regression approach, assuming reasonably high level of correlation between the two color channels. This procedure ensures more robust ratio estimation for the spots in both linear regression and traditional segmentation algorithms. The developed algorithms have been evaluated using simulated artificial images and experimental images of different designs. A demonstration version of the software can be downloaded from http://bioinfo.curie.fr/projects/maia/.

Algorithms↗

Imaginary motor movement EEG classification by Accumulative-Autocorrelation-Pulse.

Analysis of motor imaginary electroencephalogram (EEG) signals provide a feasible low-level communication channel for handicap people. We propose a classification method for imaginary right and left motor EEG using the Accumulative-Autocorrelation-Pulse (AAP) technique. This technique is based on the spatio-temporal pulse patterns generated from the accumulative autocorrelation values of selected electrodes in the ongoing EEG data. A feed forward neural network trained with the back propagation learning algorithm is used for classification. The network structure preserves and extracts the pulse-temporal feature patterns of the signal. Classification results reach 100% generalization accuracy in some single subjects and a 91% generalization over all subjects when the correct pair of electrodes are selected. Robust generalization results indicate that the autocorrelation nature of the human EEG signal contains typical patterns for classification in imaginary left and right motor events.

Adult↗

Robust detection and classification of longitudinal changes in color retinal fundus images for monitoring diabetic retinopathy.

A fully automated approach is presented for robust detection and classification of changes in longitudinal time-series of color retinal fundus images of diabetic retinopathy. The method is robust to: 1) spatial variations in illumination resulting from instrument limitations and changes both within, and between patient visits; 2) imaging artifacts such as dust particles; 3) outliers in the training data; 4) segmentation and alignment errors. Robustness to illumination variation is achieved by a novel iterative algorithm to estimate the reflectance of the retina exploiting automatically extracted segmentations of the retinal vasculature, optic disk, fovea, and pathologies. Robustness to dust artifacts is achieved by exploiting their spectral characteristics, enabling application to film-based, as well as digital imaging systems. False changes from alignment errors are minimized by subpixel accuracy registration using a 12-parameter transformation that accounts for unknown retinal curvature and camera parameters. Bayesian detection and classification algorithms are used to generate a color-coded output that is readily inspected. A multiobserver validation on 43 image pairs from 22 eyes involving nonproliferative and proliferative diabetic retinopathies, showed a 97% change detection rate, a 3% miss rate, and a 10% false alarm rate. The performance in correctly classifying the changes was 99.3%. A self-consistency metric, and an error factor were developed to measure performance over more than two periods. The average self consistency was 94% and the error factor was 0.06%. Although this study focuses on diabetic changes, the proposed techniques have broader applicability in ophthalmology.

Algorithms↗

Determining indications for care common to competing guidelines by using classification tree analysis: application to the prevention of venous thromboembolism in medical inpatients.

BACKGROUND: Substantial variations have been reported in the advice given by competing guidelines addressing the same clinical problem. OBJECTIVE: This study aimed to assess the usefulness of classification tree analysis in comparing competing guidelines. METHOD: The authors implemented a classification tree-growing algorithm on cross-sectional data from 818 patients to determine indications for prophylactic heparin treatment common to 4 competing guidelines disseminated between 1998 and 2000 and addressing the prophylaxis of venous thromboembolism in medical inpatients. RESULTS: The resulting classification tree involved 10 terminal nodes. Its mean accuracy estimated by performing 10-fold cross-validation was 82% (s=3). The guidelines consistently supported prophylactic heparin treatment for 5 indications: a previous episode of deep vein thrombosis or pulmonary embolism, recent paralysis of lower limb(s), congestive heart failure with one or more risk factors, recent myocardial infarction, and malignancy with one or more risk factors. These indications involved 257 patients (31.4%) and were supported by robust scientific evidence. Deep vein thrombosis was detected in 27 of these patients (10.5%). Two consistent negative indications involved 347 patients (42.4%). Deep vein thrombosis was detected in 9 of these patients (2.6%). Three indications involving 214 patients (26.2%) were discordant over the 4 guidelines. CONCLUSION: Classification tree analysis of real patient data is a useful strategy to identify indications common to competing guidelines. These indications should be considered for inclusion when updating guidelines. The findings of recently completed randomized trials have partly resolved the disagreement among the 4 guidelines. This approach may be helpful when developing new guidelines or for identifying topics warranting further complementary clinical trials.

Algorithms↗

Physicochemical descriptors to discriminate protein-protein interactions in permanent and transient complexes selected by means of machine learning algorithms.

Analyzing protein-protein interactions at the atomic level is critical for our understanding of the principles governing the interactions involved in protein-protein recognition. For this purpose, descriptors explaining the nature of different protein-protein complexes are desirable. In this work, the authors introduced Epic Protein Interface Classification as a framework handling the preparation, processing, and analysis of protein-protein complexes for classification with machine learning algorithms. We applied four different machine learning algorithms: Support Vector Machines, C4.5 Decision Trees, K Nearest Neighbors, and Naïve Bayes algorithm in combination with three feature selection methods, Filter (Relief F), Wrapper, and Genetic Algorithms, to extract discriminating features from the protein-protein complexes. To compare protein-protein complexes to each other, the authors represented the physicochemical characteristics of their interfaces in four different ways, using two different atomic contact vectors, DrugScore pair potential vectors and SFCscore descriptor vectors. We classified two different datasets: (A) 172 protein-protein complexes comprising 96 monomers, forming contacts enforced by the crystallographic packing environment (crystal contacts), and 76 biologically functional homodimer complexes; (B) 345 protein-protein complexes containing 147 permanent complexes and 198 transient complexes. We were able to classify up to 94.8% of the packing enforced/functional and up to 93.6% of the permanent/transient complexes correctly. Furthermore, we were able to extract relevant features from the different protein-protein complexes and introduce an approach for scoring the importance of the extracted features.

Algorithms↗

Quantifying cell scattering: the blob algorithm revisited.

BACKGROUND: A method to objectively quantify cell scattering would permit quantitative evaluation of therapies and compounds intended to affect this physiologic process, which has relevance to normal (e.g., development) and pathologic (e.g., metastasis) events. METHODS: A grid-based modified blob analysis was performed on a set of images of Madin-Darby Canine Kidney (MDCK) cells to quantify the following parameters: the number of cellular clusters in each image, the size of the clusters in terms of pixel counts, and the number of cells in each cluster. These parameters were used as measures of cell scattering and were compared with subjective assessments of scattering made by three experienced examiners. RESULTS: The quantitative parameters correlated strongly to subjective assessments. The algorithm displayed a different concept of "clustering" than the examiners and consistently identified more clusters than did the examiners. There was close agreement in the number of cells counted. All three quantitative parameters correlated strongly to the subjective scattering scores, as follows: cluster count (r(s) = -0.765 to -0.789, P < 0.0001), cluster size in pixels (r(s) = 0.838 to 0.845, P < 0.0001), and cluster size in cells (r(s) = 0.758 to 0.804, P < 0.0001). The parameters were continuous, providing greater resolving power than ordinal subjective scores. CONCLUSIONS: The findings confirmed that our algorithm reproduces the traditional classification of scattering with improved resolution, quantification, and objectivity.

Algorithms↗

Classifying antibodies using flow cytometry data: class prediction and class discovery.

Classifying monoclonal antibodies, based on the similarity of their binding to the proteins (antigens) on the surface of blood cells, is essential for progress in immunology, hematology and clinical medicine. The collaborative efforts of researchers from many countries have led to the classification of thousands of antibodies into 247 clusters of differentiation (CD). Classification is based on flow cytometry and biochemical data. In preliminary classifications of antibodies based on flow cytometry data, the object requiring classification (an antibody) is described by a set of random samples from unknown densities of fluorescence intensity. An individual sample is collected in the experiment, where a population of cells of a certain type is stained by the identical fluorescently marked replicates of the antibody of interest. Samples are collected for multiple cell types. The classification problems of interest include identifying new CDs (class discovery or unsupervised learning) and assigning new antibodies to the known CD clusters (class prediction or supervised learning). These problems have attracted limited attention from statisticians. We recommend a novel approach to the classification process in which a computer algorithm suggests to the analyst the subset of the "most appropriate" classifications of an antibody in class prediction problems or the "most similar" pairs/ groups of antibodies in class discovery problems. The suggested algorithm speeds up the analysis of a flow cytometry data by a factor 10-20. This allows the analyst to focus on the interpretation of the automatically suggested preliminary classification solutions and on planning the subsequent biochemical experiments.

Antibodies, Monoclonal↗

Classification of fMRI independent components using IC-fingerprints and support vector machine classifiers.

We present a general method for the classification of independent components (ICs) extracted from functional MRI (fMRI) data sets. The method consists of two steps. In the first step, each fMRI-IC is associated with an IC-fingerprint, i.e., a representation of the component in a multidimensional space of parameters. These parameters are post hoc estimates of global properties of the ICs and are largely independent of a specific experimental design and stimulus timing. In the second step a machine learning algorithm automatically separates the IC-fingerprints into six general classes after preliminary training performed on a small subset of expert-labeled components. We illustrate this approach in a multisubject fMRI study employing visual structure-from-motion stimuli encoding faces and control random shapes. We show that: (1) IC-fingerprints are a valuable tool for the inspection, characterization and selection of fMRI-ICs and (2) automatic classifications of fMRI-ICs in new subjects present a high correspondence with those obtained by expert visual inspection of the components. Importantly, our classification procedure highlights several neurophysiologically interesting processes. The most intriguing of which is reflected, with high intra- and inter-subject reproducibility, in one IC exhibiting a transiently task-related activation in the 'face' region of the primary sensorimotor cortex. This suggests that in addition to or as part of the mirror system, somatotopic regions of the sensorimotor cortex are involved in disambiguating the perception of a moving body part. Finally, we show that the same classification algorithm can be successfully applied, without re-training, to fMRI collected using acquisition parameters, stimulation modality and timing considerably different from those used for training.

Algorithms↗

Feature subset selection for support vector machines through discriminative function pruning analysis.

In many pattern classification applications, data are represented by high dimensional feature vectors, which induce high computational cost and reduce classification speed in the context of support vector machines (SVMs). To reduce the dimensionality of pattern representation, we develop a discriminative function pruning analysis (DFPA) feature subset selection method in the present study. The basic idea of the DFPA method is to learn the SVM discriminative function from training data using all input variables available first, and then to select feature subset through pruning analysis. In the present study, the pruning is implement using a forward selection procedure combined with a linear least square estimation algorithm, taking advantage of linear-in-the-parameter structure of the SVM discriminative function. The strength of the DFPA method is that it combines good characters of both filter and wrapper methods. Firstly, it retains the simplicity of the filter method avoiding training of a large number of SVM classifier. Secondly, it inherits the good performance of the wrapper method by taking the SVM classification algorithm into account.

Journal Article↗

Treatment of congenital facial nevi.

The treatment of congenital facial nevi is often difficult and challenging. Previous authors have reported their techniques, results, and complications when treating these lesions. Our objectives are to simplify the treatment planning by subdividing the lesions with a new classification and using this to formulate a surgical algorithm. One hundred and two patients with congenital facial nevi were reviewed. All of these patients have had surgical excision for the lesions. We have subgrouped the lesions into three groups, according to size, number of aesthetic units involved, and number of reconstructive stages required. Group I included lesions 1 to 3 cm in maximal diameter, within one aesthetic unit, and requiring one or two reconstructive stages. This group included 29 patients. Group II included lesions 3 to 12 cm in maximal diameter, covering one or two aesthetic units, and requiring not more than two stages of reconstruction. This group had 41 patients. Group III consisted of extensive lesions, over 12 cm in maximal diameter, covering several aesthetic units, and requiring several stages of reconstruction. In this group, we had 32 patients. On the basis of our experience in treating congenital facial nevi in this series, we have developed a surgical algorithm for reconstruction. We are optimistic that this will assist the surgeon in surgical planning and treating this complex patient population. The algorithm is arranged according to the new classification of congenital facial nevi that is presented.

Adolescent↗