PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Prognostic factors for survival in advanced non-small-cell lung cancer: univariate and multivariate analyses including recursive partitioning and amalgamation algorithms in 1,052 patients. The European Lung Cancer Working Party.

PURPOSE: This study attempted to determine the prognostic value for survival of various pretreatment characteristics in patients with nonresectable non-small-cell lung cancer in the context of more than 10 years of experience of a European Cooperative Group. PATIENTS AND METHODS: We included in the analysis all eligible patients (N = 1,052) with advanced non-small-cell lung cancer registered onto one of seven trials conducted by the European Lung Cancer Working Party (ELCWP) during one decade. The patients were treated by chemotherapy regimens based on platinum derivatives. We prospectively collected 23 variables and analyzed them by univariate and multivariate methods. RESULTS: The global estimated median survival time was 29 weeks, with a 95% confidence interval of 27 to 30 weeks. After univariate analysis, we applied two multivariate statistical techniques. In a Cox regression model, the selected explanatory variables were disease extent, Karnofsky performance status, WBC and neutrophil counts, metastatic involvement of skin, serum calcium level, age, and sex. These results were confirmed by application of recursive partitioning and amalgamation algorithms (RECPAM), which led to classification of the patients into four homogeneous subgroups. CONCLUSION: We confirmed by our analysis the role of well-known independent prognostic factors for survival, but also identified the effect of the neutrophil count, rarely studied, with the use of two methods: a classical Cox regression model and a RECPAM analysis. The classification of patients into the four subgroups we obtained needs to be validated in other series.

Algorithms↗

A cost-function approach to rival penalized competitive learning (RPCL).

Rival penalized competitive learning (RPCL) has been shown to be a useful tool for clustering on a set of sample data in which the number of clusters is unknown. However, the RPCL algorithm was proposed heuristically and is still in lack of a mathematical theory to describe its convergence behavior. In order to solve the convergence problem, we investigate it via a cost-function approach. By theoretical analysis, we prove that a general form of RPCL, called distance-sensitive RPCL (DSRPCL), is associated with the minimization of a cost function on the weight vectors of a competitive learning network. As a DSRPCL process decreases the cost to a local minimum, a number of weight vectors eventually fall into a hypersphere surrounding the sample data, while the other weight vectors diverge to infinity. Moreover, it is shown by the theoretical analysis and simulation experiments that if the cost reduces into the global minimum, a correct number of weight vectors is automatically selected and located around the centers of the actual clusters, respectively. Finally, we apply the DSRPCL algorithms to unsupervised color image segmentation and classification of the wine data.

Algorithms↗

Recognising vascular causes of leg complaints in endurance athletes. Part 1: validation of a decision algorithm.

Flow limitations in the iliac arteries of endurance athletes during exercise were previously ascribed solely to intravascular lesions. We postulate that functional kinking of the arteries can also result in flow limitations. However, the diagnostic tools in routine practice are not effective in diagnosing such flow limitations in a substantial proportion of athletes, mainly because these diagnostic tools do not measure in the provocative situations. Ninety-two symptomatic legs in 80 endurance athletes were examined with newly developed, sports-specific vascular tests. Thirty-five asymptomatic cyclists matched for working capacity served as the control subjects. Legs were classified as vascular or non-vascular following a decision algorithm, based upon the results of these diagnostic tests, excluding orthopaedic causes by the effects of specific treatment. Independently of this clinical classification, an alternative method was applied to find stable characteristics in the total patient group using factor analysis. This characterisation was based on scores on 14 test variables deriving from diagnostic tests that were not used in the decision algorithm, thus avoiding dependency between the clinical categorisation and the statistical categorisation. The hypothesis was that these characteristics were sufficiently sensitive to classify patients with vascular and non-vascular complaints. If so, these characteristics should correspond with the one derived from the decision algorithm. Following the decision algorithm, 58 legs (63%) were classified as vascular, 29 (32%) as non-vascular and 5 (5%) as inconclusive. The latter were considered non-vascular. In a substantial proportion of the vascular patients, kinking of the iliac arteries was identified as the major cause of flow limitation. The characteristics derived from factor analysis proved to classify 87% in agreement with the decision algorithm (kappa 0.56). The agreement is sufficient for validation of the clinical classification. The algorithm can therefore be applied in clinical situations to diagnose endurance athletes with flow limitations due to both intravascular lesions and kinking of the arteries.

Adult↗

LogitBoost classifier for discriminating thermophilic and mesophilic proteins.

A novel classifier, the so-called LogitBoost classifier, was introduced to discriminate the thermophilic and mesophilic proteins according to their primary structures. When the 20-amino acid composition was chosen as the feature vector, the overall accuracy of the self-consistency check and a five-fold cross-validation procedure was 97.0% and 86.6%, respectively. To test if the method was also applicable to a wide range of biological targets, an independent testing dataset was also used. The method based on LogitBoost algorithm has achieved an overall classification accuracy of 88.9%. According to the three different validation check approaches, it was demonstrated that LogitBoost outperformed AdaBoost and performed comparably with RBF neural network and support vector machine. The influence of protein size on discrimination was addressed.

Algorithms↗

Classification of patients on the basis of otoneurological data by using Kohonen networks.

Machine learning methods such as neural networks, decision trees and genetic algorithms can be useful to aid in the classification of patients. We tested Kohonen artificial neural networks, which are known to be effective for classification tasks. Our sample included patients with six different diseases. The Kohonen network algorithm recognized the four largest groups reliably, but the two smallest groups were too small for the method. Neural networks seem to be promising for the computer-aided classification of otoneurological patients provided that the number of patients used is sufficiently large.

Algorithms↗

Obtaining interpretable fuzzy classification rules from medical data.

For many application problems classifiers can be used to support a decision making process. In some domains-in areas like medicine especially-it is preferable not to use black box approaches. The user should be able to understand the classifier and to evaluate its results. Fuzzy rule based classifiers are especially suitable, because they consist of simple linguistically interpretable rules and do not have some of the drawbacks of symbolic or crisp rule based classifiers. Classifiers must often be created from data by a learning process, because there is not enough expert knowledge to determine their parameters completely. A simple and convenient way to learn fuzzy classifiers from data is provided by neuro-fuzzy approaches. In this paper we discuss extensions to the learning algorithms of neuro-fuzzy classification (NEFCLASS), a neuro-fuzzy approach for data analysis that we have presented before. We present interactive strategies for pruning rules and variables from a trained classifier to enhance its readability, and demonstrate our approach on a small example.

Algorithms↗

Development of clinical criteria for osteoarthritis.

Clinical criteria for the classification of osteoarthritis (OA) in the knee have been developed and the use of algorithms has been proposed. Criteria for the classification of OA in the hand and hip are still being developed.

Hand↗

Identification of carboxypeptidase E and gamma-glutamyl hydrolase as biomarkers for pulmonary neuroendocrine tumors by cDNA microarray.

Pulmonary neuroendocrine tumors vary dramatically in their malignant behavior. Their classification, based on histological examination, is often difficult. In search of molecular and prognostic markers for these tumors, we used cDNA microarray analysis of human transcripts against reference RNA from a well-characterized immortalized bronchial epithelial cell line, BEAS-2B. Tumor cells were isolated by laser-capture microdissection from primary tumors of 17 typical carcinoids, small cell lung cancers, and large cell neuroendocrine carcinomas. An unsupervised, hierarchical clustering algorithm resulted in a precise classification of each tumor subtype according to the proposed histological classification. Selection of genes, using supervised analysis, resulted in the identification of 198 statistically significant genes (P <.004) that also accurately discriminated between 3 predefined tumor subtypes. Two-by-two comparisons of these genes identified classifier genes that distinguished each tumor subtype from the others. Changes in expression of selected differentially expressed genes for each tumor subtype were internally validated by real-time reverse-transcription polymerase chain reaction. Expression of 2 potential classifier gene products, carboxypeptidase E (CPE) and gamma-glutamyl hydrolase (GGH), was validated by immunohistochemistry and cross-validated on additional archival samples of pulmonary neuroendocrine tumors. Kaplan-Meier survival analysis revealed that immunostaining for CPE was a statistically significant predictor of good prognosis, whereas GGH expression correlated with poor prognosis. Thus, cDNA microarray analysis led to the identification of 2 novel biomarkers that should facilitate molecular diagnosis and further study of pulmonary neuroendocrine tumors.

Biomarkers, Tumor↗

Non-parametric classification of protein secondary structures.

Proteins were classified into their families using a classification tree method which is based on the coefficient of variations of physico-chemical and geometrical properties of the secondary structures of proteins. The tree method uses as splitting criterion the increase in purity when a node is split into two subnodes and the size of the tree is controlled by a threshold level for the improvement of the apparent misclassification rate (AMR) of the tree after each splitting step. The classification tree method seems effective in reproducing similar structural groupings as the method of dynamic programming. For comparison, we also used another two methods: neural networks and support vector machines. We could show that the presented classification tree method performs better in classifying proteins into their families. The presented algorithm might be suitable for a rapid preliminary classification of proteins into their corresponding families.

Algorithms↗

QSAR and k-nearest neighbor classification analysis of selective cyclooxygenase-2 inhibitors using topologically-based numerical descriptors.

Experimental IC(50) data for 314 selective cyclooxygenase-2 (COX-2) inhibitors are used to develop quantitation and classification models as a potential screening mechanism for larger libraries of target compounds. Experimental log(IC(50)) values ranged from 0.23 to > or = 5.00. Numerical descriptors encoding solely topological information are calculated for all structures and are used as inputs for linear regression, computational neural network, and classification analysis routines. Evolutionary optimization algorithms are then used to search the descriptor space for information-rich subsets which minimize the rms error of a diverse training set of compounds. An eight-descriptor model was identified as a robust predictor of experimental log(IC(50)) values, producing a root-mean-square error of 0.625 log units for an external prediction set of inhibitors which took no part in model development. A k-nearest neighbor classification study of the data set discriminating between active and inactive members produced a nine-descriptor model able to accurately classify 83.3% of the prediction set compounds correctly.

Cyclooxygenase 2↗

Pocketome via comprehensive identification and classification of ligand binding envelopes.

We developed a new computational algorithm for the accurate identification of ligand binding envelopes rather than surface binding sites. We performed a large scale classification of the identified envelopes according to their shape and physicochemical properties. The predicting algorithm, called PocketFinder, uses a transformation of the Lennard-Jones potential calculated from a three-dimensional protein structure and does not require any knowledge about a potential ligand molecule. We validated this algorithm using two systematically collected data sets of ligand binding pockets from complexed (bound) and uncomplexed (apo) structures from the Protein Data Bank, 5616 and 11,510, respectively. As many as 96.8% of experimental binding sites were predicted at better than 50% overlap level. Furthermore 95.0% of the asserted sites from the apo receptors were predicted at the same level. We demonstrate that conformational differences between the apo and bound pockets do not dramatically affect the prediction results. The algorithm can be used to predict ligand binding pockets of uncharacterized protein structures, suggest new allosteric pockets, evaluate feasibility of protein-protein interaction inhibition, and prioritize molecular targets. Finally the data base of the known and predicted binding pockets for the human proteome structures, the human pocketome, was collected and classified. The pocketome can be used for rapid evaluation of possible binding partners of a given chemical compound.

Algorithms↗

Extraction subject-specific motor imagery time-frequency patterns for single trial EEG classification.

We introduce a new adaptive time-frequency plane feature extraction strategy for the segmentation and classification of electroencephalogram (EEG) corresponding to left and right hand motor imagery of a brain-computer interface task. The proposed algorithm adaptively segments the time axis by dividing the EEG data into non-uniform time segments over a dyadic tree. This is followed by grouping the expansion coefficients in the frequency axis in each segment. The most discriminative features are selected from the segmented time-frequency plane and fed to a linear discriminant for classification. The proposed algorithm achieved an average classification accuracy of 84.3% on six subjects by selecting the most discriminant subspaces for each one. For comparison, classification results based on an autoregressive model are also presented where the mean accuracy of the same subjects turned out to be 79.5%. Interestingly the subjects and two hemispheres of each subject are represented by distinct segmentations and features. This indicates that the proposed method can handle inter-subject variability when constructing brain-computer interfaces.

Algorithms↗

Spatial partitioning using multivariate cluster analysis and a contiguity algorithm.

Spatial analysis of epidemiological data can be a useful tool for identifying patterns of disease occurrence and can provide substantial support for prevention and control strategies. To obtain the greatest spatial resolution, it is important to use the smallest available areal units with homogeneous population. However, small areas usually have a small population, introducing spurious variability in the chosen indicators of disease occurrence. This paper describes an approach for combining small geographical units to stabilize mortality rates by pooling information across areas according to specified risk profiles. The procedure is based on a principal component analysis, followed by a cluster analysis of social-economic indicators to classify the risk profile of each small area. The classification is used in an algorithm to join neighbouring areas with similar profiles until an estimated population size is achieved. We applied this method to two Administrative Regions of the city of Rio de Janeiro, Brazil, using the census tracts as the basic areal unit. Census tracts were classified according to four socioeconomic categories distributed spatially as a mosaic, where tracts of differing categories neighbour each other. The aggregation algorithm produced a new partition of the region studied, with the created areal units preserving the internal socioeconomic homogeneity.

Adult↗

Patellar dislocation in army conscripts.

Between 1990 and 1996, 119 Finnish male conscripts underwent operative treatment for patellar dislocation. There were 68 conscripts (58%) with primary and 51 conscripts (42%) with recurrent patellar dislocation. Sixty-five (55%) dislocations occurred during military service, 40 (34%) occurred during sports activities, and 14 (11%) occurred during leisure time. The most common cause of injury was military training at battle exercises (n = 30). The typical injury mechanism was knee valgus rotation on fixed foot and tibia (97 conscripts, 82%). Surgical procedures performed were open in 75 conscripts (63%) and arthroscopically assisted in 44 conscripts (37%). Twenty-three (19%) redislocations occurred during follow-up (mean, 6 years; range, 3-9 years). The subjective outcome of treatment was excellent in 23 (19%), good in 42 (35%), moderate in 44 (37%), and poor in 10 (9%) conscripts. The most common residual complaint was patellofemoral pain (25 conscripts, 21%). Only 42 (35%) conscripts were able to finish their military service normally; fitness classification was decreased in 16 conscripts (13%), and 61 (52%) were temporarily exempted from military service (class E). The results of operative treatment did not differ significantly in conscripts with primary and recurrent dislocation, except for the time of first recurrence, which was significantly longer in conscripts with primary than with recurrent dislocation (27 versus 9 months). Patellar dislocation is the most common form of severe knee injury among conscripts and significantly hampers military service. Operative treatment yields only satisfactory results. Preventive measurements should be considered. A suggested algorithm for the treatment and classification of acute and recurrent patellar dislocation among military conscripts is presented.

Acute Disease↗

Novel statistical classification model of type 2 diabetes mellitus patients for tailor-made prevention using data mining algorithm.

To estimate the usefulness of data mining algorithms for extracting risk predictors of diabetic vascular complications in proper order in the future, we tried applying the Classification and Regression Trees (CART) method to the prevalence data of 165 type 2 diabetic outpatients and already known risk factors. Among the 6 categorical and 15 continuous risk factors, age (cutoff: 65.4) was the best predictor for classifying patients into groups with and without macroangiopathy (p=0.000). Body weight (cutoff: 53.9) was the best predictor (p=0.006) in the older group (age >65.4), whereas systolic blood pressure (cutoff: 144.5) was the best predictor in the remaining group (p=0.002). Age (cutoff: 64.8) was also the best predictor for categorizing them into groups with and without microangiopathy (p=0.000). In the older group (age >64.8), BMI (cutoff: 21.5) was the best predictor (p=0.001), whereas morbidity term (cutoff: 15.5) was the best predictor in the other group (p=0.01 0). Because the orders and values of all risk factors and cutoff points mined were reasonable clinically, this method may have the potential to highlight predictors in order of importance to apply tailor-made prevention of diabetic vascular complications.

Algorithms↗

Molecular classification of breast cancer patients by gene expression profiling.

For many tumors, pathological subclasses exist which have to be further defined by genetic markers to improve therapy and follow-up strategies. In this study, cDNA array analyses of breast cancers have been performed to classify tumors into categories based on expression patterns. Comparing purified normal ductal epithelial cells and corresponding tumour tissues, the expression of only a small fraction of genes was found to be significantly changed. A subset of genes repeatedly found to be differentially expressed in breast cancers was subsequently employed to perform a classification of 82 normal and malignant breast specimens by cluster analysis. This analysis identifies a subgroup of transcriptionally related tumours, designated class A, which can be further subdivided into A1 and A2. Correlation with classical clinicopathological parameters revealed that subgroup A1 was characterized by a high number of node-positive tumours (14 of 16). In this subgroup there was a disproportionate number of patients who had already developed distant metastases at the time of diagnosis (25% in this subgroup, compared with 5% among the rest of the samples). Taken together, the use of these differentially expressed marker genes in conjunction with sample clustering algorithms provides a novel molecular classification of breast cancer specimens, which facilitates the identification of patients with a higher risk of recurrence.

Breast Neoplasms↗

Sequence database search using jumping alignments.

We describe a new algorithm for amino acid sequence classification and the detection of remote homologues. The rationale is to exploit both vertical and horizontal information of a multiple alignment in a well balanced manner. This is in contrast to established methods like profiles and hidden Markov models which focus on vertical information as they model the columns of the alignment independently. In our setting, we want to select from a given database of "candidate sequences" those proteins that belong to a given superfamily. In order to do so, each candidate sequence is separately tested against a multiple alignment of the known members of the superfamily by means of a new jumping alignment algorithm. This algorithm is an extension of the Smith-Waterman algorithm and computes a local alignment of a single sequence and a multiple alignment. In contrast to traditional methods, however, this alignment is not based on a summary of the individual columns of the multiple alignment. Rather, the candidate sequence at each position is aligned to one sequence of the multiple alignment, called the "reference sequence". In addition, the reference sequence may change within the alignment, while each such jump is penalized. To evaluate the discriminative quality of the jumping alignment algorithm, we compared it to hidden Markov models on a subset of the SCOP database of protein domains. The discriminative quality was assessed by counting the number of false positives that ranked higher than the first true positive (FP-count). For moderate FP-counts above five, the number of successful searches with our method was considerably higher than with hidden Markov models.

Animals↗

Utilization of artificial neural networks in the diagnosis of optic nerve diseases.

This research is concentrated on the diagnosis of optic nerve disease through the analysis of pattern electroretinography (PERG) signals with the help of artificial neural network (ANN). Multilayer feed forward ANN trained with a Levenberg Marquart (LM) backpropagation algorithm was implemented. The designed classification structure has about 96.4% sensitivity, 90.4% specifity and positive prediction is calculated to be 94.2%. The end results are classified as healthy and diseased. Testing results were found to be compliant with the expected results that are derived from the physician's direct diagnosis. The end benefit would be to assist the physician to make the final decision without hesitation.

Adult↗