PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

QSAR and classification study of 1,4-dihydropyridine calcium channel antagonists based on least squares support vector machines.

The least squares support vector machine (LSSVM), as a novel machine learning algorithm, was used to develop quantitative and classification models as a potential screening mechanism for a novel series of 1,4-dihydropyridine calcium channel antagonists for the first time. Each compound was represented by calculated structural descriptors that encode constitutional, topological, geometrical, electrostatic, quantum-chemical features. The heuristic method was then used to search the descriptor space and select the descriptors responsible for activity. Quantitative modeling results in a nonlinear, seven-descriptor model based on LSSVM with mean-square errors 0.2593, a predicted correlation coefficient (R(2)) 0.8696, and a cross-validated correlation coefficient (R(cv)(2)) 0.8167. The best classification results are found using LSSVM: the percentage (%) of correct prediction based on leave one out cross-validation was 91.1%. This paper provides a new and effective method for drug design and screening.

Algorithms↗

A new classifier based on information theoretic learning with unlabeled data.

Supervised learning is conventionally performed with pairwise input-output labeled data. After the training procedure, the adaptive system's weights are fixed while the testing procedure with unlabeled data is performed. Recently, in an attempt to improve classification performance unlabeled data has been exploited in the machine learning community. In this paper, we present an information theoretic learning (ITL) approach based on density divergence minimization to obtain an extended training algorithm using unlabeled data during the testing. The method uses a boosting-like algorithm with an ITL based cost function. Preliminary simulations suggest that the method has the potential to improve the performance of classifiers in the application phase.

Algorithms↗

Cancer gene search with data-mining and genetic algorithms.

Cancer leads to approximately 25% of all mortalities, making it the second leading cause of death in the United States. Early and accurate detection of cancer is critical to the well being of patients. Analysis of gene expression data leads to cancer identification and classification, which will facilitate proper treatment selection and drug development. Gene expression data sets for ovarian, prostate, and lung cancer were analyzed in this research. An integrated gene-search algorithm for genetic expression data analysis was proposed. This integrated algorithm involves a genetic algorithm and correlation-based heuristics for data preprocessing (on partitioned data sets) and data mining (decision tree and support vector machines algorithms) for making predictions. Knowledge derived by the proposed algorithm has high classification accuracy with the ability to identify the most significant genes. Bagging and stacking algorithms were applied to further enhance the classification accuracy. The results were compared with that reported in the literature. Mapping of genotype information to the phenotype parameters will ultimately reduce the cost and complexity of cancer detection and classification.

Algorithms↗

Virtual screening of molecular databases using a support vector machine.

The Support Vector Machine (SVM) is an algorithm that derives a model used for the classification of data into two categories and which has good generalization properties. This study applies the SVM algorithm to the problem of virtual screening for molecules with a desired activity. In contrast to typical applications of the SVM, we emphasize not classification but enrichment of actives by using a modified version of the standard SVM function to rank molecules. The method employs a simple and novel criterion for picking molecular descriptors and uses cross-validation to select SVM parameters. The resulting method is more effective at enriching for active compounds with novel chemistries than binary fingerprint-based methods such as binary kernel discrimination.

Algorithms↗

A fast learning algorithm for deep belief nets.

We show how to use "complementary priors" to eliminate the explaining-away effects that make inference difficult in densely connected belief nets that have many hidden layers. Using complementary priors, we derive a fast, greedy algorithm that can learn deep, directed belief networks one layer at a time, provided the top two layers form an undirected associative memory. The fast, greedy algorithm is used to initialize a slower learning procedure that fine-tunes the weights using a contrastive version of the wake-sleep algorithm. After fine-tuning, a network with three hidden layers forms a very good generative model of the joint distribution of handwritten digit images and their labels. This generative model gives better digit classification than the best discriminative learning algorithms. The low-dimensional manifolds on which the digits lie are modeled by long ravines in the free-energy landscape of the top-level associative memory, and it is easy to explore these ravines by using the directed connections to display what the associative memory has in mind.

Algorithms↗

Effectiveness of environmental cluster analysis in representing regional species diversity.

A major challenge of regional conservation planning is the identification of sets of sites that together represent the overall biodiversity of the relevant region. Environmental cluster analysis (ECA) has been proposed as a potential tool for efficient selection of conservation sites, but the consequences of methodological decisions involved in its application have not been tested so far. We evaluated the performance of ECA with respect to two such decisions: the choice of the clustering algorithm (single linkage, complete linkage, unweighted arithmetic average, unweighted centroid, Ward's minimum variance, and the ALOC algorithm) and the weight given to different groups of environmental variables (rainfall, temperature, and lithology). Specifically we tested how these decisions affect the spatial configuration of clusters of sites defined by the ECA, whether and how they affect the effectiveness of the ECA (i.e., its ability to represent regional species diversity), and whether the effectiveness of alternative methods of hierarchical clustering can be predicted a priori based on the cophenetic correlation. We used an extensive database of the flora of Israel to test these questions. Differences in both the clustering algorithm and the weighting regime had considerable effects on the spatial configuration of the ECA clusters. The single-linkage algorithm produced mostly single-cell clusters plus a single large-sized cluster and was therefore found inappropriate for environmental regionalization. The effectiveness of the ECA was also sensitive to changes in the clustering algorithm and the weighting regime. Yet, most combinations of clustering algorithms and weighting regimes performed significantly better in capturing regional biodiversity than random null models. The main deviation was classifications based on Ward's minimum variance algorithm, which performed less well relative to all other algorithms. The two algorithms that showed the highest effectiveness (unweighted average and unweighted centroid clustering) also exhibited the highest values of the cophenetic correlation, suggesting that this index may serve as a potential indicator for the effectiveness of alternative ECA algorithms.

Algorithms↗

Efficient training of RBF networks for classification.

Radial Basis Function networks with linear outputs are often used in regression problems because they can be substantially faster to train than Multi-layer Perceptrons. For classification problems, the use of linear outputs is less appropriate as the outputs are not guaranteed to represent probabilities. We show how RBFs with logistic and softmax outputs can be trained efficiently using the Fisher scoring algorithm. This approach can be used with any model which consists of a generalised linear output function applied to a model which is linear in its parameters. We compare this approach with standard non-linear optimisation algorithms on a number of datasets.

Algorithms↗

Sequence-based source tracking of Escherichia coli based on genetic diversity of beta-glucuronidase.

High levels of fecal bacteria are a concern for recreational waters; however, the source of contamination is often unknown. This study investigated whether direct sequencing of a bacterial gene could be utilized for detecting genetic differences between bacterial strains for microbial source tracking. A 525-nucleotide segment of the gene for beta-glucuronidase (uidA) was sequenced in 941 Escherichia coli isolates from the Clinton River-Lake St. Clair watershed, 182 E. coli isolates from human and animal feces, and 34 E. coli isolates from a combined sewer. Environmental isolates exhibited 114 alleles in 11 groups on a genetic tree. Frequency of strains from different genetic groups differed significantly (p < 0.03) between upstream reaches (Bear Creek-Red Run), downstream reaches, and Lake St. Clair beaches. Fecal E. coli uidA sequences exhibited 81 alleles that overlapped with the environmental set. An algorithm to assign alleles to different host sources averaged approximately 75% correct classification with the fecal data set. Using the same algorithm, the percent of environmental isolates assignable to humans decreased significantly between Bear Creek-Red Run (30 +/- 3%) and the beaches (17 +/- 2%) (p < 0.05). Birds accounted for approximately 50% of assignable environmental isolates. For combined sewer isolates, the same algorithm assigned 51% to humans. These experiments demonstrate differences in the frequency of different E. coli strains at different locations in a watershed, and provide a "proof in principle" that sequence-based data can be used for microbial source tracking.

DNA, Bacterial↗

Fine-grained structural classification of biosynthetic gene cluster-encoded products.

MOTIVATION: Biosynthetic gene clusters (BGCs) are responsible the biosynthesis of many natural products, including a multitude of effective therapeutics and their precursors. Advances in genomic data collection as well as computational techniques have made it possible to identify BGCs at scale. However, accurately determining the types of BGC-encoded products from genomic content remains elusive. RESULTS: Here, we introduce BGC annotation tool (BGCat), a machine learning method for fine-grained structural classification of BGC-encoded products, leveraging the NPClassifier natural product nomenclature. Our method leverages a pre-trained protein language model for creating meaningful gene representations and a deep neural network for class label prediction. We show the method outperforms state-of-the-art approaches in coarse-grained product classification and is effective for detailed classification. We implement a clustering-based augmentation strategy for BGC-product relationships, addressing a crucial gap in the available datasets. We then introduce the concept of product class profiles of gene cluster families (GCFs), associating each GCF with a probabilistic distribution of product types and offering a new perspective on GCF functions. Lastly, we use BGCat to provide new product class labels for over 100k BGCs in antiSMASH DB that presently have minimal information about their products. AVAILABILITY AND IMPLEMENTATION: The source code and trained model weights are freely available at https://github.com/HassounLab/BGCat.

Multigene Family↗

Raman spectroscopy for diagnosis of atherosclerosis: a rapid analysis using neural networks.

Near-infrared Raman spectroscopy (NIRS) is one of the novel techniques that has a potential for in vivo diagnosis of atherosclerosis in human arteries. For such real time clinical applications, a rapid collection and analysis of the data is needed. One of the major problems with the fast data collection is that the noise generated by the detector has the same level as the Raman signal from the tissue, which makes the analysis difficult. In this work, NIRS measurements have been carried out on a total of 60 samples from human coronary arteries. Raman spectral data with the correlated histopathological analysis have been used as a basis to stimulate the cases of severe noise conditions. The main objective of this paper is the comparison of different processing algorithms that have been developed based on either wavelet transformation or principal component analysis for compressing the Raman spectral vectors and a rapid data classification based on different neural network architectures. The developed algorithms found to provide promising diagnosis results with classification errors smaller than 5%, even in the cases of Raman data with collection times as small as 20 ms. It has been concluded that the developed algorithms would be very much useful in the development of Raman spectroscopy systems for in vivo biological applications.

Algorithms↗

Spectral mapping tools from the earth sciences applied to spectral microscopy data.

BACKGROUND: Spectral imaging, originating from the field of earth remote sensing, is a powerful tool that is being increasingly used in a wide variety of applications for material identification. Several workers have used techniques like linear spectral unmixing (LSU) to discriminate materials in images derived from spectral microscopy. However, many spectral analysis algorithms rely on assumptions that are often violated in microscopy applications. This study explores algorithms originally developed as improvements on early earth imaging techniques that can be easily translated for use with spectral microscopy. METHODS: To best demonstrate the application of earth remote sensing spectral analysis tools to spectral microscopy data, earth imaging software was used to analyze data acquired with a Leica confocal microscope with mechanical spectral scanning. For this study, spectral training signatures (often referred to as endmembers) were selected with the ENVI (ITT Visual Information Solutions, Boulder, CO) "spectral hourglass" processing flow, a series of tools that use the spectrally over-determined nature of hyperspectral data to find the most spectrally pure (or spectrally unique) pixels within the data set. This set of endmember signatures was then used in the full range of mapping algorithms available in ENVI to determine locations, and in some cases subpixel abundances of endmembers. RESULTS: Mapping and abundance images showed a broad agreement between the spectral analysis algorithms, supported through visual assessment of output classification images and through statistical analysis of the distribution of pixels within each endmember class. CONCLUSIONS: The powerful spectral analysis algorithms available in COTS software, the result of decades of research in earth imaging, are easily translated to new sources of spectral data. Although the scale between earth imagery and spectral microscopy is radically different, the problem is the same: mapping material locations and abundances based on unique spectral signatures.

Algorithms↗

Identifying spatial relationships in neural processing using a multiple classification approach.

The application of statistical classification methods to in vivo functional neuroimaging data makes it possible to explore spatial patterns in task-related changes in neural processing. Cluster analysis is one group of descriptive statistical procedures that can assist in identifying classes of brain regions that exhibit similar task-related functionality. In practice, a limitation of cluster analysis is that the performances of clustering algorithms rely on unknown characteristics of the data, making it difficult to determine which procedure best suits a particular analysis. We present a multiple classification approach that incorporates numerous algorithms, evaluates the associated classifications, and either selects a plausible partition relative to the others considered or pools the results from the numerous methods. The multiple classification approach utilizes a new performance criterion, called the relative information (RI) measure, to evaluate the quality of the candidate partitions and as the basis for producing a composite classification image. Employing multiple classifications, rather than a single algorithm, our methodology increases the chance of detecting the functional relationships within the data and, therefore, produces more reliable results. We apply our methodology to a PET study to explore spatial relationships in measured brain function associated with increasing blood alcohol concentration levels, and we perform a simulation study to evaluate the performance of RI.

Alcoholic Intoxication↗

The BCI Competition 2003: progress and perspectives in detection and discrimination of EEG single trials.

Interest in developing a new method of man-to-machine communication--a brain-computer interface (BCI)--has grown steadily over the past few decades. BCIs create a new communication channel between the brain and an output device by bypassing conventional motor output pathways of nerves and muscles. These systems use signals recorded from the scalp, the surface of the cortex, or from inside the brain to enable users to control a variety of applications including simple word-processing software and orthotics. BCI technology could therefore provide a new communication and control option for individuals who cannot otherwise express their wishes to the outside world. Signal processing and classification methods are essential tools in the development of improved BCI technology. We organized the BCI Competition 2003 to evaluate the current state of the art of these tools. Four laboratories well versed in EEG-based BCI research provided six data sets in a documented format. We made these data sets (i.e., labeled training sets and unlabeled test sets) and their descriptions available on the Internet. The goal in the competition was to maximize the performance measure for the test labels. Researchers worldwide tested their algorithms and competed for the best classification results. This paper describes the six data sets and the results and function of the most successful algorithms.

Adult↗

Local structural motifs of protein backbones are classified by self-organizing neural networks.

Important and relevant information is expected to be encoded in local structural elements of proteins. An unsupervised learning algorithm (Kohonen algorithm) was applied to the representation and unbiased classification of local backbone structures contained in a set of proteins. Training yielded a two-dimensional Kohonen feature map with 100 different structural motifs including certain helical and strand structures. All motifs were represented in a phi-psi-plot and some of them as a three-dimensional model. The course of structural motifs along the backbone of four selected proteins (cytochrome b5, cytochrome b562, lysozyme, gamma crystallin) was investigated in detail. Trajectories and histograms visualizing the abundance of characteristic motifs allowed for the distinction between different types of protein overall folds. It is demonstrated how the histograms may be used to construct a structural similarity matrix for proteins. The Kohonen algorithm provides a simple procedure for classification of local protein structures independent of any a priori knowledge of leading structural motifs. Training of the Kohonen network leads to the generation of "consensus structures' serving for the task of classification.

Algorithms↗

Simulating patients with Parallel Health State Networks.

The American Board of Family Practice is developing a computer-based recertification process to generate patient simulations from a knowledge base. Simulated patients require a stochastically generated history and response to treatment, suggesting a Monte Carlo-like patient generation process. Knowledge acquisition experiments revealed that description of a patient's overall health as a node in a Monte Carlo model was difficult for domain experts to use, severely limited knowledge reusability, and created a plethora of awkwardly defined health states. We explored a model in which patients traverse several parallel health state networks simultaneously, so that overall health is a vector describing the current nodes from every Parallel Network. This model has a reasonable biological basis, more easily defined data, and greatly improved reuse potential, at the cost of more complex simulation algorithms. Experiments using osteoarthritis stages, weight classification, and absence or presence of gastric ulcers as three Parallel Networks demonstrate the feasibility of this approach to simulating patients.

Algorithms↗

Chronic fatigue syndrome--a clinically empirical approach to its definition and study.

BACKGROUND: The lack of standardized criteria for defining chronic fatigue syndrome (CFS) has constrained research. The objective of this study was to apply the 1994 CFS criteria by standardized reproducible criteria. METHODS: This population-based case control study enrolled 227 adults identified from the population of Wichita with: (1) CFS (n = 58); (2) non-fatigued controls matched to CFS on sex, race, age and body mass index (n = 55); (3) persons with medically unexplained fatigue not CFS, which we term ISF (n = 59); (4) CFS accompanied by melancholic depression (n = 27); and (5) ISF plus melancholic depression (n = 28). Participants were admitted to a hospital for two days and underwent medical history and physical examination, the Diagnostic Interview Schedule, and laboratory testing to identify medical and psychiatric conditions exclusionary for CFS. Illness classification at the time of the clinical study utilized two algorithms: (1) the same criteria as in the surveillance study; (2) a standardized clinically empirical algorithm based on quantitative assessment of the major domains of CFS (impairment, fatigue, and accompanying symptoms). RESULTS: One hundred and sixty-four participants had no exclusionary conditions at the time of this study. Clinically empirical classification identified 43 subjects as CFS, 57 as ISF, and 64 as not ill. There was minimal association between the empirical classification and classification by the surveillance criteria. Subjects empirically classified as CFS had significantly worse impairment (evaluated by the SF-36), more severe fatigue (documented by the multidimensional fatigue inventory), more frequent and severe accompanying symptoms than those with ISF, who in turn had significantly worse scores than the not ill; this was not true for classification by the surveillance algorithm. CONCLUSION: The empirical definition includes all aspects of CFS specified in the 1994 case definition and identifies persons with CFS in a precise manner that can be readily reproduced by both investigators and clinicians.

Adult↗

Modeling drug detection and diagnosis with the 'drug evaluation and classification program'.

In this study, we propose formal models and algorithms to detect drug impairment and identify the impairing drug type, on the basis of data obtained by a Drug Evaluation and Classification (DEC) investigation. The DEC program relies on measurements of vital signs and observable signs and symptoms. A formal model, based on data collected by police officers trained to detect and identify drug impairments, yielded sensitivity levels greater than 60% and specificity levels greater than 90% for impairments caused by cannabis, alprazolam, and amphetamine. For codeine, with a specificity of nearly 90% the sensitivity was only 20%. Using logistic regression, the formal model was much more accurate than the trained officers in identifying impairments from cannabis, alprazolam, and amphetamine. Both the formal model and the officers were quite poor in identifying codeine impairment. In conclusion, the joint application of the DECP procedures with the formal model is useful for drug detection and identification.

Accidents, Traffic↗