PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Radial plication in concentric mastopexy.

Concentric mastopexy presents many challenges to the plastic surgeon, especially when breast augmentation is part of the treatment plan. Radial plication is a reproducible and accurate technique for elevating the nipple-areolar complex and shaping the breast mound. Patient selection is important to the success of the radial plication procedure and concentric mastopexy in general. Although most surgeons agree that patients with smaller degrees of nipple ptosis and smaller breasts have better results than patients with greater degrees of nipple ptosis and larger breasts, there has never been an algorithm for patient selection. Regnault's classification of breast ptosis addresses the degree of nipple ptosis, but no consideration is given to breast volume. Radial placation proved to be a valuable tool in the treatment of 87 patients undergoing concentric mastopexy in the author's practice over the past 30 months. An algorithm addressing degrees of breast ptosis and breast volume is provided. The plastic surgeon can anticipate gratifying results if the algorithm provided is incorporated into his or her patient selection for concentric mastopexy. The concentric mastopexy technique is similar to the tailor tack procedure for standard mastopexy, allowing the plastic surgeon to mold and shape the breast before making a critical incision.

Breast↗

Supervised learning-based cell image segmentation for p53 immunohistochemistry.

In this paper, we present two new algorithms for cell image segmentation. First, we demonstrate that pixel classification-based color image segmentation in color space is equivalent to performing segmentation on grayscale image through thresholding. Based on this result, we develop a supervised learning-based two-step procedure for color cell image segmentation, where color image is first mapped to grayscale via a transform learned through supervised learning, thresholding is then performed on the grayscale image to segment objects out of background. Experimental results show that the supervised learning-based two-step procedure achieved a boundary disagreement (mean absolute distance) of 0.85 while the disagreement produced by the pixel classification-based color image segmentation method is 3.59. Second, we develop a new marker detection algorithm for watershed-based separation of overlapping or touching cells. The merit of the new algorithm is that it employs both photometric and shape information and combines the two naturally in the framework of pattern classification to provide more reliable markers. Extensive experiments show that the new marker detection algorithm achieved 0.4% and 0.2% over-segmentation and under-segmentation, respectively, while reconstruction-based method produced 4.4% and 1.1% over-segmentation and under-segmentation, respectively.

Algorithms↗

Phylogenetic classification of human papillomaviruses: correlation with clinical manifestations.

Human papillomaviruses (HPVs) are a heterogeneous group of small dsDNA viruses which cause a variety of proliferative epithelial lesions at specific anatomical sites. Although more than 65 different virus types have been cloned and characterized, no uniform classification system exists. In order to classify HPV DNA types, phylogenetic trees were constructed based on nucleotide sequence alignments using parsimony and distance matrix algorithms. The resulting phylogenetic trees provide a classification of the HPVs into specific groups encompassing the known tissue tropism and oncogenic potential of each HPV type. The implications of a phylogenetic taxonomy on the diagnostic detection of HPVs and the concept of different HPV species are discussed.

Algorithms↗

A decomposition model to track gene expression signatures: preview on observer-independent classification of ovarian cancer.

MOTIVATION: A number of algorithms and analytical models have been employed to reduce the multidimensional complexity of DNA array data and attempt to extract some meaningful interpretation of the results. These include clustering, principal components analysis, self-organizing maps, and support vector machine analysis. Each method assumes an implicit model for the data, many of which separate genes into distinct clusters defined by similar expression profiles in the samples tested. A point of concern is that many genes may be involved in a number of distinct behaviours, and should therefore be modelled to fit into as many separate clusters as detected in the multidimensional gene expression space. The analysis of gene expression data using a decomposition model that is independent of the observer involved would be highly beneficial to improve standard and reproducible classification of clinical and research samples. RESULTS: We present a variational independent component analysis (ICA) method for reducing high dimensional DNA array data to a smaller set of latent variables, each associated with a gene signature. We present the results of applying the method to data from an ovarian cancer study, revealing a number of tissue type-specific and tissue type-independent gene signatures present in varying amounts among the samples surveyed. The observer independent results of such molecular analysis of biological samples could help identify patients who would benefit from different treatment strategies. We further explore the application of the model to similar high-throughput studies.

Algorithms↗

Real-time learning capability of neural networks.

In some practical applications of neural networks, fast response to external events within an extremely short time is highly demanded and expected. However, the extensively used gradient-descent-based learning algorithms obviously cannot satisfy the real-time learning needs in many applications, especially for large-scale applications and/or when higher generalization performance is required. Based on Huang's constructive network model, this paper proposes a simple learning algorithm capable of real-time learning which can automatically select appropriate values of neural quantizers and analytically determine the parameters (weights and bias) of the network at one time only. The performance of the proposed algorithm has been systematically investigated on a large batch of benchmark real-world regression and classification problems. The experimental results demonstrate that our algorithm can not only produce good generalization performance but also have real-time learning and prediction capability. Thus, it may provide an alternative approach for the practical applications of neural networks where real-time learning and prediction implementation is required.

Computer Systems↗

Evolutionary algorithms for multiobjective and multimodal optimization of diagnostic schemes.

This paper addresses the optimization of noninvasive diagnostic schemes using evolutionary algorithms in medical applications based on the interpretation of biosignals. A general diagnostic methodology using a set of definable characteristics extracted from the biosignal source followed by the specific diagnostic scheme is presented. In this framework, multiobjective evolutionary algorithms are used to meet not only classification accuracy but also other objectives of medical interest, which can be conflicting. Furthermore, the use of both multimodal and multiobjective evolutionary optimization algorithms provides the medical specialist with different alternatives for configuring the diagnostic scheme. Some application examples of this methodology are described in the diagnosis of a specific cardiac disorder-paroxysmal atrial fibrillation.

Algorithms↗

Advances in thoracostomy tube management.

This article summarizes several of the studies utilizing randomized trials or predetermined algorithms for chest tube management. The classification system, when to use wall suction, when to use water seal, and how to safely discharge patients by the fourth postoperative day-even with air leaks-are outlined.

Algorithms↗

A statistical framework for the classification of tensor morphologies in diffusion tensor images.

Tractography algorithms for diffusion tensor (DT) images consecutively connect directions of maximal diffusion across neighboring DTs in order to reconstruct the 3-dimensional trajectories of white matter tracts in vivo in the human brain. The performance of these algorithms, however, is strongly influenced by the amount of noise in the images and by the presence of degenerate tensors-- i.e., tensors in which the direction of maximal diffusion is poorly defined. We propose a simple procedure for the classification of tensor morphologies that uses test statistics based on invariant measures of DTs, such as fractional anisotropy, while accounting for the effects of noise on tensor estimates. Examining DT images from seven human subjects, we demonstrate that this procedure validly classifies DTs at each voxel into standard types (nondegenerate DTs, as well as degenerate oblate, prolate or isotropic DTs), and we provide preliminary estimates for the prevalence and spatial distribution of degenerate tensors in these brains. We also show that the P values for test statistics are more sensitive tools for classifying tensor morphologies than are invariant measures of anisotropy alone.

Algorithms↗

Using classification tree and logistic regression methods to diagnose myocardial infarction.

Early and accurate diagnosis of myocardial infarction (MI) in patients who present to the Emergency Room (ER) complaining of chest pain is an important problem in emergency medicine. A number of decision aids have been developed to assist with this problem but have not achieved general use. Machine learning techniques, including classification tree and logistic regression (LR) methods, have the potential to create simple but accurate decision aids. Both a classification tree (FT Tree) and an LR model (FT LR) have been developed to predict the probability that a patient with chest pain is having an MI based solely upon data available at time of presentation to the ER. Training data came from a data set collected in Edinburgh, Scotland. Each model was then tested on a separate Edinburgh data set, as well as on a data set from a different hospital in Sheffield, England. Previously published models, the Goldman classification tree[1] and Kennedy LR equation[2], were evaluated on the same test data sets. On the Edinburgh test set, results showed that the FT Tree, FT LR, and Kennedy LR performed equally well, with ROC curve areas of 94.04%, 94.28%, and 94.30%, respectively, while the Goldman Tree's performance was significantly poorer, with an area of 84.03%. The difference in ROC areas between the first three models and the Goldman model is significant beyond the 0.0001 level. On the Sheffield test set, results showed that the FT Tree, FT LR, and Kennedy LR ROC areas were not significantly different (p > = 0.17), while the FT Tree again outperformed the Goldman Tree (p = 0.006). Unlike previous work[3], this study indicates that classification trees, which have certain advantages over LR models, may perform as well as LR models in the diagnosis of patients with MI.

Algorithms↗

[The proliferative activity of myelokaryocytes and the cellular composition of the bone marrow].

Flow cytometry was used to study myelokaryocyte distribution according to the stage of the cellular cycle in 167 bone marrow specimens 94 of which were taken by puncture from hemoblastosis and anemia patients. The results obtained were compared with myelogram data. It has been established that the method provides stable and reliable values, irrespective of cellular composition of the puncture specimens. Basing on the recurrent algorithm of J. H. Fridman's classification, a computer program has been derived that permitted differential diagnosis to be made based on the data of cytometry and myelogram.

Algorithms↗

Stability-based validation of clustering solutions.

Data clustering describes a set of frequently employed techniques in exploratory data analysis to extract "natural" group structure in data. Such groupings need to be validated to separate the signal in the data from spurious structure. In this context, finding an appropriate number of clusters is a particularly important model selection question. We introduce a measure of cluster stability to assess the validity of a cluster model. This stability measure quantifies the reproducibility of clustering solutions on a second sample, and it can be interpreted as a classification risk with regard to class labels produced by a clustering algorithm. The preferred number of clusters is determined by minimizing this classification risk as a function of the number of clusters. Convincing results are achieved on simulated as well as gene expression data sets. Comparisons to other methods demonstrate the competitive performance of our method and its suitability as a general validation tool for clustering solutions in real-world problems.

Algorithms↗

A new criterion to classify globular proteins based on their secondary structure contents.

MOTIVATION: With the enlargement of protein structure databases, it is hoped that a method to classify proteins automatically will be developed. Although the classification criterion proposed by Nakashima et al. ( J. Biochem., 1986, 99, 153-162) was widely used in the literature, it leads to some inconsistencies with the classification databases currently available in the class assignment of protein structures. To improve their work, a new classification criterion is proposed relying on statistical analysis of the secondary structure contents of more than 200 proteins with well-known structural classes. The Fisher linear discriminant algorithm is used to derive the new classification criterion. RESULTS: Three cross-validation tests are performed to evaluate the new criterion. In the jackknife test, of the 210 proteins used to derive the criterion, 206 are correctly classified with an accuracy of 98.10%. Of the 16 proteins of purely intermediate structure (i.e. structures lying near borderlines between two classes) in the first test set, 15 are correctly classified with an accuracy of 93.75%. For the second test set which consists of 200 proteins selected randomly from SCOP, a testing accuracy of 94.00% is obtained. For comparison, the criterion of Nakashima et al. is also used to classify the 210, 16 and 200 proteins, respectively. Consequently, accuracies of 94.76%, 62.50% and 91.50% are obtained, respectively. On average, the accuracy of the new classification criterion is 4% higher than that of Nakashima et al. AVAILABILITY: The program is available on request from the first author. CONTACT: ctzhang@tju.edu.cn

Algorithms↗

Searching protein sequence libraries: comparison of the sensitivity and selectivity of the Smith-Waterman and FASTA algorithms.

The sensitivity and selectivity of the FASTA and the Smith-Waterman protein sequence comparison algorithms were evaluated using the superfamily classification provided in the National Biomedical Research Foundation/Protein Identification Resource (PIR) protein sequence database. Sequences from each of the 34 superfamilies in the PIR database with 20 or more members were compared against the protein sequence database. The similarity scores of the related and unrelated sequences were determined using either the FASTA program or the Smith-Waterman local similarity algorithm. These two sets of similarity scores were used to evaluate the ability of the two comparison algorithms to identify distantly related protein sequences. The FASTA program using the ktup = 2 sensitivity setting performed as well as the Smith-Waterman algorithm for 19 of the 34 superfamilies. Increasing the sensitivity by setting ktup = 1 allowed FASTA to perform as well as Smith-Waterman on an additional 7 superfamilies. The rigorous Smith-Waterman method performed better than FASTA with ktup = 1 on 8 superfamilies, including the globins, immunoglobulin variable regions, calmodulins, and plastocyanins. Several strategies for improving the sensitivity of FASTA were examined. The greatest improvement in sensitivity was achieved by optimizing a band around the best initial region found for every library sequence. For every superfamily except the globins and immunoglobulin variable regions, this strategy was as sensitive as a full Smith-Waterman. For some sequences, additional sensitivity was achieved by including conserved but nonidentical residues in the lookup table used to identify the initial region.

Algorithms↗

A new variational shape-from-orientation approach to correcting intensity inhomogeneities in magnetic resonance images.

A new intensity inhomogeneity correction algorithm based on a variational shape-from-orientation formulation is presented. Unlike most previous methods, the proposed algorithm is fully automatic, widely applicable and very efficient. Since no prior classification knowledge about the image is assumed in the proposed algorithm, it can be applied to correct intensity inhomogeneities for a wide variety of medical images. In this paper, a finite-element method is used to model the smooth bias-field function. Orientation constraints for the bias-field function are computed at the nodal locations of the regular discretization grid away from the boundary between different class regions. The selection of reliable orientation constraints is facilitated by the goodness of fit of a first-order polynomial model to the neighborhood of each nodal location. The automatically selected orientation constraints are integrated in a regularization framework, which leads to minimization of a convex and quadratic energy function. This energy minimization is accomplished by solving a linear system with a large, sparse, symmetric and positive semi-definite stiffness matrix. We employ an adaptive preconditioned conjugate-gradient algorithm to solve the linear system very efficiently. Experimental results on a variety of magnetic resonance images are given to demonstrate the effectiveness and efficiency of the proposed algorithm.

Algorithms↗

Biclustering in gene expression data by tendency.

The advent of DNA microarray technologies has revolutionized the experimental study of gene expression. Clustering is the most popular approach of analyzing gene expression data and has indeed proven to be successful in many applications. Our work focuses on discovering a subset of genes which exhibit similar expression patterns along a subset of conditions in the gene expression matrix. Specifically, we are looking for the Order Preserving clusters (OPCluster), in each of which a subset of genes induce a similar linear ordering along a subset of conditions. The pioneering work of the OPSM model[3], which enforces the strict order shared by the genes in a cluster, is included in our model as a special case. Our model is more robust than OPSM because similarly expressed conditions are allowed to form order equivalent groups and no restriction is placed on the order within a group. Guided by our model, we design and implement a deterministic algorithm, namely OPCTree, to discover OP-Clusters. Experimental study on two real datasets demonstrates the effectiveness of the algorithm in the application of tissue classification and cell cycle identification. In addition, a large percentage of OP-Clusters exhibit significant enrichment of one or more function categories, which implies that OP-Clusters indeed carry significant biological relevance.

Algorithms↗

Difficulties in diagnosing hypertension: implications and alternatives.

OBJECTIVE: To estimate the magnitude of misclassification rates with commonly used algorithms for the detection of hypertensives and to suggest a sequential approach to screening. DESIGN: A conventional statistical model was used with several different algorithms to determine the number and types of errors made in categorizing two different populations, a general population sample and a population with a high risk of hypertension. METHODS: The calculations were made for single-visit screens, similar to those used in epidemiologic studies, for three-visit screens commonly used in clinical practice and clinical trials for cutoff points of 85, 95 and 105 mmHg. A sequential probability ratio screen was proposed and the error rates estimated. RESULTS: Perhaps only one-third to two-thirds of people whose measured diastolic pressures exceed 95 mmHg actually have average pressures that high. The disparity between a single measured diastolic pressure and the mean of many pressure values also leads to errors in identifying individual subjects with mild hypertension. In a general population, single measurements of diastolic pressure exceed 95 mmHg in approximately equal numbers of normotensive, borderline and hypertensive subjects; moreover, one-third of those who are usually in the hypertensive range are not identified. All commonly used screening algorithms give too many false-positive and/or false-negative results. A sequential screening algorithm averaged 3.8 visits per subject and identified 95% of the hypertensives, with only 2.5% of those identified having usual diastolic pressures below 90 mmHg. CONCLUSIONS: Population-based surveys like the National Health and Nutrition Examination Survey (NHANES) may markedly overestimate the true prevalence of hypertension. This overestimate is greatest for mild hypertension and could significantly affect the cost/benefit analyses of public health policy. Alternative screening methods, such as the sequential algorithm proposed, may have significant benefits in providing a correct classification.

Adult↗

Medical linguistics: automated indexing into SNOMED.

This paper reviews the state of the art in processing medical language data. The area is divided into the topics: (1) morphologic analysis, (2) syntactic analysis, (3) semantic analysis, and (4) pragmatics. Additional attention is given to medical nomenclatures and classifications as the bases of (automated) indexing procedures which are required whenever medical information is formalized. These topics are completed by an evaluation of related data structures and methods used to organize language-based medical knowledge.

Abstracting and Indexing↗

ICF and ICD codes provide a standard language of disability in young children.

BACKGROUND AND OBJECTIVES: The aim of this study was to examine the utility of a hierarchical algorithm incorporating codes from the International Classification of Functioning, Disability and Health--ICF (WHO, 2001) and the International Statistical Classification of Diseases-ICD (WHO, 1994) to classify reasons for eligibility of young children in early intervention. METHODS: The database for this study was a nationally representative enrollment sample of more than 5,500 children in a longitudinal study of early intervention. Reasons for eligibility were reviewed and matched to the closest ICF or ICD codes under one of four major categories (Body Functions/Structures, Activities/Participation, Health Conditions, and Environmental Factors). RESULTS: The average number of reasons for eligibility provided per child was 1.5, resulting in a population summary exceeding 100%. A total of 305 ICF and ICD codes were used with most (77%) of the children having codes in the category of Body Function/Structures. Forty-one percent of the sample had codes of Health Conditions, whereas the proportions with codes in the Activities/Partipication and Environmental Categories were 10 and 5%, respectively. CONCLUSIONS: The results demonstrate that ICD and ICF can be jointly used as a common language to document disability characteristics of children in early intervention.

Activities of Daily Living↗