PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

A new technique for identifying sequence heterochrony.

Sequence heterochrony (changes in the order in which events occur) is a potentially important, but relatively poorly explored, mechanism for the evolution of development. In part, this is because of the inherent difficulties in inferring sequence heterochrony across species. The event-pairing method, developed independently by several workers in the mid-1990s, encodes sequences in a way that allows them to be examined in a phylogenetic framework, but the results can be difficult to interpret in terms of actual heterochronic changes. Here, we describe a new, parsimony-based method to interpret such results. For each branch of the tree, it identifies the least number of event movements (heterochronies) that will explain all the observed event-pair changes. It has the potential to find all alternative, equally parsimonious explanations, and generate a consensus, containing the movements that form part of every equally most parsimonious explanation. This new technique, which we call Parsimov, greatly increases the utility of the event-pair method for inferring instances of sequence heterochrony.

Algorithms↗

SDM: a fast distance-based approach for (super) tree building in phylogenomics.

Phylogenomic studies aim to build phylogenies from large sets of homologous genes. Such "genome-sized" data require fast methods, because of the typically large numbers of taxa examined. In this framework, distance-based methods are useful for exploratory studies and building a starting tree to be refined by a more powerful maximum likelihood (ML) approach. However, estimating evolutionary distances directly from concatenated genes gives poor topological signal as genes evolve at different rates. We propose a novel method, named super distance matrix (SDM), which follows the same line as average consensus supertree (ACS; Lapointe and Cucumel, 1997) and combines the evolutionary distances obtained from each gene into a single distance supermatrix to be analyzed using a standard distance-based algorithm. SDM deforms the source matrices, without modifying their topological message, to bring them as close as possible to each other; these deformed matrices are then averaged to obtain the distance supermatrix. We show that this problem is equivalent to the minimization of a least-squares criterion subject to linear constraints. This problem has a unique solution which is obtained by resolving a linear system. As this system is sparse, its practical resolution requires O(naka) time, where n is the number of taxa, k the number of matrices, and a < 2, which allows the distance supermatrix to be quickly obtained. Several uses of SDM are proposed, from fast exploratory studies to more accurate approaches requiring heavier computing time. Using simulations, we show that SDM is a relevant alternative to the standard matrix representation with parsimony (MRP) method, notably when the taxa sets of the different genes have low overlap. We also show that SDM can be used to build an excellent starting tree for an ML approach, which both reduces the computing time and increases the topogical accuracy. We use SDM to analyze the data set of Gatesy et al. (2002, Syst. Biol. 51: 652-664) that involves 48 genes of 75 placental mammals. The results indicate that these genes have strong rate heterogeneity and confirm the simulation conclusions.

Algorithms↗

Phylogenetic diversity within seconds.

We consider a (phylogenetic) tree with n labeled leaves, the taxa, and a length for each branch in the tree. For any subset of k taxa, the phylogenetic diversity is defined as the sum of the branch-lengths of the minimal subtree connecting the taxa in the subset. We introduce two time-efficient algorithms (greedy and pruning) to compute a subset of size k with maximal phylogenetic diversity in O(n log k) and O[n + (n-k) log (n-k)] time, respectively. The greedy algorithm is an efficient implementation of the so-called greedy strategy (Steel, 2005; Pardi and Goldman, 2005), whereas the pruning algorithm provides an alternative description of the same problem. Both algorithms compute within seconds a subtree with maximal phylogenetic diversity for trees with 100,000 taxa or more.

Algorithms↗

Childhood leukaemia relapse risk factors. A rough sets approach.

A rough sets approach was applied to a data set consisting of clinical and laboratory examinations (condition attributes) of children with acute lymphoblastic leukaemia to generate a set of rules for the prediction of disease relapse (conclusion attributes). The information system is presented as a table composed of 69 rows corresponding to the patients and 16 columns corresponding to the attributes. Using manipulation based on rough set theory the information system is reduced to get a subset of a minimum number of attributes ensuring an acceptable quality of classification. Then the conclusion algorithm derived from the reduced system is presented as a conclusion table. The relationship between condition and conclusion attributes is being shown. The research leads to the conclusion that intensive, high dose central nervous system prophylactic irradiation seems to be a better prevention against CNS relapse. Rough set theory is a useful and still complementary tool of medical (biological) data analysis.

Adolescent↗

A Hypercard program for the identification of biological specimens.

A Hypercard-based software tool developed to provide help in the identification of biological specimens is presented. The package implements a matching algorithm that compares alphanumeric strings and runs on Macintosh computers, though its simple architecture can be transferred to other computers and/or other programming environments. The overall performance of the program and its easy customization to specific problems are demonstrated by discussing at length one application in the field of earthworm identification.

Algorithms↗

PCP: a program for supervised classification of gene expression profiles.

UNLABELLED: PCP (Pattern Classification Program) is an open-source machine learning program for supervised classification of patterns (vectors of measurements). The principal use of PCP in bioinformatics is design and evaluation of classifiers for use in clinical diagnostic tests based on measurements of gene expression. PCP implements leading pattern classification and gene selection algorithms and incorporates cross-validation estimation of classifier performance. Importantly, the implementation integrates gene selection and class prediction stages, which is vital for computing reliable performance estimates in small-sample scenarios. Additionally, the program includes automated and efficient model selection (optimization of parameters) for support vector machine (SVM) classifier. The distribution includes Linux and Windows/Cygwin binaries. The program can easily be ported to other platforms. AVAILABILITY: Free download at http://pcp.sourceforge.net

Algorithms↗

Inferring phylogenetic networks by the maximum parsimony criterion: a case study.

Horizontal gene transfer (HGT) may result in genes whose evolutionary histories disagree with each other, as well as with the species tree. In this case, reconciling the species and gene trees results in a network of relationships, known as the "phylogenetic network" of the set of species. A phylogenetic network that incorporates HGT consists of an underlying species tree that captures vertical inheritance and a set of edges which model the "horizontal" transfer of genetic material. In a series of papers, Nakhleh and colleagues have recently formulated a maximum parsimony (MP) criterion for phylogenetic networks, provided an array of computationally efficient algorithms and heuristics for computing it, and demonstrated its plausibility on simulated data. In this article, we study the performance and robustness of this criterion on biological data. Our findings indicate that MP is very promising when its application is extended to the domain of phylogenetic network reconstruction and HGT detection. In all cases we investigated, the MP criterion detected the correct number of HGT events required to map the evolutionary history of a gene data set onto the species phylogeny. Furthermore, our results indicate that the criterion is robust with respect to both incomplete taxon sampling and the use of different site substitution matrices. Finally, our results show that the MP criterion is very promising in detecting HGT in chimeric genes, whose evolutionary histories are a mix of vertical and horizontal evolution. Besides the performance analysis of MP, our findings offer new insights into the evolution of 4 biological data sets and new possible explanations of HGT scenarios in their evolutionary history.

Algorithms↗

Characterization of the integrity of three-dimensional trabecular bone microstructure by connectivity and shape analysis using high-resolution magnetic resonance imaging in vivo.

Bone mineral density and bone structure are the main determinants of bone strength in osteoporosis. In this study we used high-resolution magnetic resonance imaging to visualize the bone microstructure in the finger phalanges in vivo and to assess the topological three-dimensional connectivity of the trabecular network and the shape of the trabeculae as measures of bone quality. We visualized the phalanges of young and elderly healthy volunteers in vivo with a spatial resolution of 152 microm x 152 microm x 280 microm. Image processing software to quantify three measures of connectedness was developed and tested: connectivity, global connectivity density, and local connectivity density. Global three-dimensional connectivity ranged from 904 to 1,607 connections. Global connectivity density ranged from 2.9 to 4.7 connections per mm with large intersubject differences. We found a decrease of local connectivity density with growing distance from the joint ranging from 5.1 to 0.2 connections per mm. These preliminary results represent a quantitative description of the well-known rarefication of the trabecular network when moving from epiphysis to the diaphysis. Three-dimensional visualization showed a dense network consisting mostly of rod-like trabeculae at the epiphysis changing to a less dense network of a few plate-like structures near the medullary canal. An algorithm for the quantitative classification of trabecular architecture with regard to plate or rod-like shape was tested for feasibility. We conclude that in vivo assessment of three-dimensional properties of the trabecular network is possible in human phalanges. Determination of connectivity and shape will allow quantification of structural aspects of osteoporotic changes and may improve assessment of fracture risk.

Adult↗

Utility scores for the Health Utilities Index Mark 2: an empirical assessment of alternative mapping functions.

INTRODUCTION: The Health Utilities Index is one of the most widely used generic health status classification systems. The valuation algorithm rests upon a power transformation between visual analog scale (VAS) and standard gamble (SG) data. This transformation has been the subject of much debate. To date, the literature has concentrated upon the mapping functions themselves. We examine whether alternative mapping functions produce more accurate utility predictions. METHODS: We undertook valuation interviews with 201 members of the UK general population, following the methods of the original Health Utilities Index-2 valuation survey. We estimated a cubic and a power mapping function using the mean VAS and SG data from the survey and calculated 2 alternative Multiplicative Multi Attribute Utility Functions (MAUFs). Using a validation sample, we assessed the predictive precision of the models in terms of accuracy (root mean square error and mean absolute error); clinical importance of the prediction error (% states with prediction error greater than 0.03); bias (t test); and whether the prediction error was related to the health state severity (Ljung Box Q statistic). RESULTS: The power MAUF was an extremely poor predictive model, mean absolute error = 0.18, root mean square error = 0.206. The predictions were biased (t = -12.92). The errors were not related to the severity of the health state, (Liung Box = 10.87). The Cubic MAUF was a better predictive model than the Power MAUF (mean absolute error = 0.084, root mean square error = 0.101). The Cubic MAUF also produced biased predictions (t = -3.57). The prediction errors were not related to the severity of the health state (Liung Box = 5.242). DISCUSSION: The Power MAUF is considerably worse than the Cubic MAUF. Our results suggest that the problems with the power function can translate into significant problems with predictive performance of the MAUF.

Activities of Daily Living↗

Current approach to radial nerve paralysis.

LEARNING OBJECTIVES: After studying this article, the participant should be able to: 1. Identify all potential points of radial nerve compression and other likely causes of radial nerve injury. 2. Accurately diagnose both surgical and nonsurgical causes of radial nerve paralysis. 3. Define a safe and effective approach to the surgical release and reconstruction of the radial nerve. Radial nerve paralysis, which can result from a complex humerus fracture, direct nerve trauma, compressive neuropathies, neuritis, or (rarely) from malignant tumor formation, has been reported throughout the literature, with some controversy regarding its diagnosis and management. The appropriate management of any radial nerve palsy depends primarily on an accurate determination of its cause, severity, duration, and level of involvement. The radial nerve can be injured as proximally as the brachial plexus or as distally as the posterior interosseous or radial sensory nerve. This article reviews the etiology, prognosis, and various treatments available for radial nerve paralysis. It also provides a new classification system and treatment algorithm to assist in the management of patients with radial nerve palsies, and it offers a simple, five-step approach to radial nerve release in the forearm.

Algorithms↗

Characterization of renal allograft rejection by urinary proteomic analysis.

OBJECTIVE: To develop a diagnostic method with no morbidity or mortality for the detection of acute renal transplant rejection. SUMMARY BACKGROUND DATA: Rejection constitutes the major impediment to the success of transplantation. Currently available methods, including clinical presentation and biochemical organ function parameters, often fail to detect rejection until late stages of progression. Renal biopsies have associated morbidity and mortality and provide only a limited sample of the organ. METHODS: Thirty-four urine samples were collected from 32 renal transplant patients at various stages posttransplantation. Samples were collected from 17 transplant recipients with acute rejection and 15 patients with no rejection. Samples from patients less than 4 days posttransplant were omitted from data analysis due to the presence of excessive inflammatory response proteins. Rejection status was confirmed by kidney biopsy. Specimens were analyzed in triplicate using SELDI mass spectrometry. The obtained spectra were subjected to bioinformatic analysis using ProPeak as well as CART (Classification and Regression Tree) algorithms to identify rejection biomarker candidates. These candidates were identified by their molecular weight and ranked by their ability to distinguish between nonrejection and rejection based on receiver operating characteristic (ROC) analysis. The candidates with the highest area under the ROC curve (AUC) exhibited the best diagnostic performance. RESULTS: The best candidate biomarkers demonstrated highly successful diagnostic performance: 6.5 kd (AUC = 0.839, P <.0001), 6.7 kd (AUC = 0.839, P <.0001), 6.6 kd (AUC = 0.807, P <.0001), 7.1 kd (AUC = 0.807, P <.0001), and 13.4 kd (AUC = 0.804, P <.0001). A separate analysis using the CART algorithm in the Ciphergen Biomarker Pattern Software correctly classified 91% of the 34 specimens in the training set, giving a sensitivity of 83% and specificity of 100% using two separate biomarker candidates at 10.0 kd and 3.4 kd. CONCLUSIONS: Biomarker candidates exist in urine that have the ability to distinguish between renal transplant patients with no rejection and those with acute rejection. These biomarker candidates are the basis for development of a noninvasive method of diagnosing acute rejection without the morbidity and mortality associated with needle biopsy. The combination of biomarkers into a panel for diagnosis leads to the possibility of enhanced diagnostic performance.

Acute Disease↗

Paired MEG data set source localization using recursively applied and projected (RAP) MUSIC.

An important class of experiments in functional brain mapping involves collecting pairs of data corresponding to separate "Task" and "Control" conditions. The data are then analyzed to determine what activity occurs during the Task experiment but not in the Control. Here we describe a new method for processing paired magnetoencephalographic (MEG) data sets using our recursively applied and projected multiple signal classification (RAP-MUSIC) algorithm. In this method the signal subspace of the Task data is projected against the orthogonal complement of the Control data signal subspace to obtain a subspace which describes spatial activity unique to the Task. A RAP-MUSIC localization search is then performed on this projected data to localize the sources which are active in the Task but not in the Control data. In addition to dipolar sources, effective blocking of more complex sources, e.g., multiple synchronously activated dipoles or synchronously activated distributed source activity, is possible since these topographies are well-described by the Control data signal subspace. Unlike previously published methods, the proposed method is shown to be effective in situations where the time series associated with Control and Task activity possess significant cross correlation. The method also allows for straightforward determination of the estimated time series of the localized target sources. A multiepoch MEG simulation and a phantom experiment are presented to demonstrate the ability of this method to successfully identify sources and their time series in the Task data.

Algorithms↗

Tracheobronchomalacia and excessive dynamic airway collapse.

Tracheobronchomalacia and excessive dynamic airway collapse are two separate forms of dynamic central airway obstruction that may or may not coexist. These entities are increasingly recognized as asthma and COPD imitators. The understanding of these disease processes, however, has been compromised over the years because of uncertainties regarding their definitions, pathogenesis and aetiology. To date, there is no standardized classification, diagnosis or management algorithm. In this article we comprehensively review the aetiology, morphopathology, physiology, diagnosis and treatment of these entities.

Airway Obstruction↗

Serum protein profiles to identify head and neck cancer.

PURPOSE: New and more consistent biomarkers of head and neck squamous cell carcinoma (HNSCC) are needed to improve early detection of disease and to monitor successful patient management. The purpose of this study was to determine whether a new proteomic technology could correctly identify protein expression profiles for cancer in patient serum samples. EXPERIMENTAL DESIGN: Surface-enhanced laser desorption/ionization-time of flight-mass spectrometry ProteinChip system was used to screen for differentially expressed proteins in serum from 99 patients with HNSCC and 102 normal controls. Protein peak clustering and classification analyses of the surface-enhanced laser desorption/ionization spectral data were performed using the Biomarker Wizard and Biomarker Patterns software (version 3.0), respectively (Ciphergen Biosystems, Fremont, CA). RESULTS: Several proteins, with masses ranging from 2778 to 20800 Da, were differentially expressed between HNSCC and the healthy controls. The serum protein expression profiles were used to develop and train a classification and regression tree algorithm, which reliably achieved a sensitivity of 83.3% and a specificity of 100% in discriminating HNSCC from normal controls. CONCLUSIONS: We propose that this technique has potential for the development of a screening test for the detection of HNSCC.

Adult↗

Slit scan flow cytometry of isolated chromosomes following fluorescence hybridization: an approach of online screening for specific chromosomes and chromosome translocations.

The recently developed methods of non radioactive in situ hybridization of chromosomes offer new aspects for chromosome analysis. Fluorescent labelling of hybridized chromosomes or chromosomal subregions allows to facilitate considerably the detection of specific chromosomal abnormalities. For many biomedical applications (e.g. biological dosimetry in the low dose range), a fast scoring for aberrations (e.g. dicentrics or translocations) in required. Here, we present an approach depending on fluorescence in situ hybridization of isolated suspension chromosomes that indicates the feasibility of a rapid screening for specific chromosomes or translocations by slit scan flow cytometry. Chromosomes of a Chinese hamster x human hybrid cell line were hybridized in suspension with biotinylated human genomic DNA. This DNA was decorated with FITC by a double antibody system against biotin. For flow cytometry the chromosomes were stabilized with ethanol and counterstained with DAPI or propidium iodide (PI). An experimental data set of several hundred double profiles was obtained by two parameter slit scan flow cytometry and evaluated automatically. The evaluation algorithm developed allowed a classification of chromosomes according to the number of centromeres and their chromosomal positions in less than 1 msec per individual profile. Approximately 20% of the measured DAPI profiles showed a bimodal distribution with a significant centromeric dip indicating a "normal" chromosomal morphology and a correct alignment in the flow system. In many cases, profiles of a "normal" bimodal fluorescence distribution of the DNA stain (DAPI, PI) were correlated with a "normal" FITC profile. Due to their centromeric indices these profiles agreed well to the expected human chromosomes of the cell line. In some cases of "normal" DAPI (PI) profiles, "aberrant" FITC profiles were observed.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

ECG processing techniques based on neural networks and bidirectional associative memories.

Two ECG processing techniques are described for the classification of QRSs, PVCs and normal and ischaemic beats. The techniques use neural network (NN) technology in two ways. The first technique, uses nonlinear ECG mapping preprocessing and subsequently for classification uses a shrinking algorithm based on NNs. This technique is applied to the QRS/PVC problem with good result. The second technique is based on the Bidirectional Associative Memory (BAM) NN and is used to distinguish normal from ischaemic beats. In this technique the ECG beat is treated as a digitized image which is then transformed into a bipolar vector suitable for input in the BAM. The results show that this method, if properly calibrated, can result in a fast and reliable ischaemic beat detection algorithm.

Algorithms↗

Automated electrocardiogram analysis: the state of the art.

The overwhelming number of electrocardiograms (ECGs) now recorded routinely has prompted the development of computer analysis which in turn has benefited electrocardiography with important technological advances. All automated ECG analysis systems adopt a similar approach: a measurement program and a program that interprets the clinical significance of these measurements along with a rhythm analysis algorithm. Measurement, selection and classification of parameters vary according to the program used. Data compression is applied to the signal to reduce processing time and allow long-term storage. Diagnostic accuracy, however, is not greatly improved over that of experienced cardiologists. Programs studied using a validated data bank provided by an international group of cardiologists show a variability not only in parameter measurement but also in diagnostic statement and in the way in which such statements are expressed. Recommendations for measurement standards have been made to fulfil the need for exchange of diagnostic criteria. No recommendations concerning the selection of parameters have been proposed, and so new parameters or combinations of parameters can be introduced with the ultimate aim of increased diagnostic performance.

Diagnosis, Computer-Assisted↗