PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Minimizing stochastic complexity using local search and GLA with applications to classification of bacteria.

In this paper, we compare the performance of two iterative clustering methods when applied to an extensive data set describing strains of the bacterial family Enterobacteriaceae. In both methods, the classification (i.e. the number of classes and the partitioning) is determined by minimizing stochastic complexity. The first method performs the minimization by repeated application of the generalized Lloyd algorithm (GLA). The second method uses an optimization technique known as local search (LS). The method modifies the current solution by making global changes to the class structure and it, then, performs local fine-tuning to find a local optimum. It is observed that if we fix the number of classes, the LS finds a classification with a lower stochastic complexity value than GLA. In addition, the variance of the solutions is much smaller for the LS due to its more systematic method of searching. Overall, the two algorithms produce similar classifications but they merge certain natural classes with microbiological relevance in different ways.

Algorithms↗

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans↗

Nonlinear fisher discriminant analysis using a minimum squared error cost function and the orthogonal least squares algorithm.

The nonlinear discriminant function obtained using a minimum squared error cost function can be shown to be directly related to the nonlinear Fisher discriminant (NFD). With the squared error cost function, the orthogonal least squares (OLS) algorithm can be used to find a parsimonious description of the nonlinear discriminant function. Two simple classification techniques will be introduced and tested on a number of real and artificial data sets. The results show that the new classification technique can often perform favourably compared with other state of the art classification techniques.

Algorithms↗

Framewise phoneme classification with bidirectional LSTM and other neural network architectures.

In this paper, we present bidirectional Long Short Term Memory (LSTM) networks, and a modified, full gradient version of the LSTM learning algorithm. We evaluate Bidirectional LSTM (BLSTM) and several other network architectures on the benchmark task of framewise phoneme classification, using the TIMIT database. Our main findings are that bidirectional networks outperform unidirectional ones, and Long Short Term Memory (LSTM) is much faster and also more accurate than both standard Recurrent Neural Nets (RNNs) and time-windowed Multilayer Perceptrons (MLPs). Our results support the view that contextual information is crucial to speech processing, and suggest that BLSTM is an effective architecture with which to exploit it.

Algorithms↗

Protein subcellular location prediction.

The function of a protein is closely correlated with its subcellular location. With the rapid increase in new protein sequences entering into data banks, we are confronted with a challenge: is it possible to utilize a bioinformatic approach to help expedite the determination of protein subcellular locations? To explore this problem, proteins were classified, according to their subcellular locations, into the following 12 groups: (1) chloroplast, (2) cytoplasm, (3) cytoskeleton, (4) endoplasmic reticulum, (5) extracell, (6) Golgi apparatus, (7) lysosome, (8) mitochondria, (9) nucleus, (10) peroxisome, (11) plasma membrane and (12) vacuole. Based on the classification scheme that has covered almost all the organelles and subcellular compartments in an animal or plant cell, a covariant discriminant algorithm was proposed to predict the subcellular location of a query protein according to its amino acid composition. Results obtained through self-consistency, jackknife and independent dataset tests indicated that the rates of correct prediction by the current algorithm are significantly higher than those by the existing methods. It is anticipated that the classification scheme and concept and also the prediction algorithm can expedite the functionality determination of new proteins, which can also be of use in the prioritization of genes and proteins identified by genomic efforts as potential molecular targets for drug design.

Algorithms↗

Structure analysis and classification of cervical cells using a processing system based on TV.

This paper presents preliminary results of a cell classification experiment using a new approach for feature extraction. The algorithm takes into account the special requirements of a fast parallel processing system (processor-oriented algorithms). A cell image is described by several hundred features derived from the nucleus only. The most significant features with respect to classification are determined by statistical analysis. Applying principal axis transform, a new feature set is computed, reduced considerably in dimensions. The data base (1,925 cell images of Papanicolaou-stained cervical specimens) was divided into a training set (963 images) and a test set (962 images). The classification results of the test set show that the recognition rate for the two-class problem (normal, suspicious) is better than 91%, using only ten morphologic features.

Cervix Mucus↗

Assessing swine thermal comfort by image analysis of postural behaviors.

Postural behavior is an integral response of animals to complex environmental factors. Huddling, nearly contacting one another on the side, and spreading are common postural behaviors of group-housed animals undergoing cold, comfortable, and warm/hot sensations, respectively. These postural patterns have been routinely used by animal caretakers to assess thermal comfort of the animals and to make according adjustment on the environmental settings or management schemes. This manual adjustment approach, however, has the inherent limitations of daily discontinuity and inconsistency between caretakers in interpretation of the animal comfort behavior. The goal of this project was to explore a novel, automated image analysis system that would assess the thermal comfort of swine and make proper environmental adjustments to enhance animal wellbeing and production efficiency. This paper describes the progress and on-going work toward the achievement of our proposed goal. The feasibility of classifying the thermal comfort state of young pigs by neural network (NN) analysis of their postural images was first examined. It included exploration of using certain feature selections of the postural behavioral images as the input to a three-layer NN that was trained to classify the corresponding thermal comfort state as being cold, comfortable, or warm. The image feature selections, a critical step for the classification, examined in this study included Fourier coefficient (FC), moment (M), perimeter and area (P&A), and combination of M and P&A of the processed binary postural images. The result was positive, with the combination of M and P&A as the input feature to the NN yielding the highest correct classification rate. Subsequent work included the development of hardware and computational algorithms that enable automatic image segmentation, motion detection, and the selection of the behavioral images suitable for use in the classification. Work is in progress to quantify the relationships of postural behavior and physiological responses of pigs using thermographs. The results are expected to facilitate objective training of NN, hence improving the accuracy of the postural image-based assessment of the thermal comfort state. Work is also in progress to implement the analysis and assessment algorithms into computer codes for real-time application.

Animal Welfare↗

New methods for the analysis of binarized BIOLOG GN data of Vibrio species: minimization of stochastic complexity and cumulative classification.

We apply minimization of stochastic complexity and the closely related method of cumulative classification to analyse the extensively studied BIOLOG GN data of Vibrio spp. Minimization of stochastic complexity provides an objective tool of bacterial taxonomy as it produces classifications that are optimal from the point of view of information theory. We compare the outcome of our results with previously published classifications of the same data set. Our results both confirm earlier detected relationships between species and discover new ones.

Algorithms↗

[Systemic sclerosis - diagnosis and classification].

Systemic sclerosis (SSc) is a polymorphic and heterogenic systemic disorder with inflammation, fibrosis and vascular damage. Early diagnosis and classification may be difficult if disease expression is oligosymptomatic (undifferentiated), presenting with only Raynaud's phenomenon or limited scleroderma. Scleroderma specific antinuclear autoantibodies, which are present early and persistently in about 90% of the patients with SSc, play an important taxonomic role. Scleroderma specific findings in nailfold capillary microscopy are sensitive and predictive for evolving SSc. An algorithm will be presented for the diagnosis and classification of SSc using clinical, capillaroscopic and serologic criteria, which are also useful for mixed or special forms of SSc. The 6th Outcome Measures in Rheumatology Clinical Trials (OMERACT) conference proposed different outcome measurements for clinical studies, however, for daily clinical practice there is as yet no consensus on status indices for disease activity, disease related damage or suitable prognostic criteria.

Algorithms↗

Comparing stage of change measures in adolescent smokers.

This study compared stage of change measures among adolescent smokers. Participants were 56 adolescents who had received smoking tickets. They filled out an assessment packet including readiness to change measures [i.e., algorithm, Crittenden's measure, University of Rhode Island Change Assessment (URICA), and Ladder], smoking history, a 30-day calendar, Fagerstrom, self-efficacy measures, locus of control, and a problem screen. The Crittenden algorithm was correlated with the Ladder and URICA motivation scores as well as with decreased frequency/amount of smoking and past quit attempts. Ladder scores were correlated with less cigarettes smoked, self-efficacy, and fewer adolescent problems. Readiness to change was unrelated to nicotine dependence and locus of control. The classification of participants into stages by the Crittenden algorithm was associated with a significant MANOVA (Wilks' Lambda F=1.56, P<.025), with group differences on reported quit attempts, abstinence self-efficacy, and adolescent problems. The Crittenden measure and the URICA motivation score were sensitive to a MET intervention.

Adolescent↗

Survival analysis with time-varying regression effects using a tree-based approach.

Nonproportional hazards often arise in survival analysis, as is evident in the data from the International Non-Hodgkin's Lymphoma Prognostic Factors Project. A tree-based method to handle such survival data is developed for the assessment and estimation of time-dependent regression effects under a Cox-type model. The tree method approximates the time-varying regression effects as piecewise constants and is designed to estimate change points in the regression parameters. A fast algorithm that relies on maximized score statistics is used in recursive segmentation of the time axis. Following the segmentation, a pruning algorithm with optimal properties similar to those of classification and regression trees (CART) is used to determine a sparse segmentation. Bootstrap resampling is used in correcting for overoptimism due to split point optimization. The piecewise constant model is often more suitable for clinical interpretation of the regression parameters than the more flexible spline models. The utility of the algorithm is shown on the lymphoma data, where we further develop the published International Risk Index into a time-varying risk index for non-Hodgkin's lymphoma.

Algorithms↗

Brain-computer interfaces for 1-D and 2-D cursor control: designs using volitional control of the EEG spectrum or steady-state visual evoked potentials.

We have developed and tested two electroencephalogram (EEG)-based brain-computer interfaces (BCI) for users to control a cursor on a computer display. Our system uses an adaptive algorithm, based on kernel partial least squares classification (KPLS), to associate patterns in multichannel EEG frequency spectra with cursor controls. Our first BCI, Target Practice, is a system for one-dimensional device control, in which participants use biofeedback to learn voluntary control of their EEG spectra. Target Practice uses a KPLS classifier to map power spectra of 62-electrode EEG signals to rightward or leftward position of a moving cursor on a computer display. Three subjects learned to control motion of a cursor on a video display in multiple blocks of 60 trials over periods of up to six weeks. The best subject's average skill in correct selection of the cursor direction grew from 58% to 88% after 13 training sessions. Target Practice also implements online control of two artifact sources: 1) removal of ocular artifact by linear subtraction of wavelet-smoothed vertical and horizontal electrooculograms (EOG) signals, 2) control of muscle artifact by inhibition of BCI training during periods of relatively high power in the 40-64 Hz band. The second BCI, Think Pointer, is a system for two-dimensional cursor control. Steady-state visual evoked potentials (SSVEP) are triggered by four flickering checkerboard stimuli located in narrow strips at each edge of the display. The user attends to one of the four beacons to initiate motion in the desired direction. The SSVEP signals are recorded from 12 electrodes located over the occipital region. A KPLS classifier is individually calibrated to map multichannel frequency bands of the SSVEP signals to right-left or up-down motion of a cursor on a computer display. The display stops moving when the user attends to a central fixation point. As for Target Practice, Think Pointer also implements wavelet-based online removal of ocular artifact; however, in Think Pointer muscle artifact is controlled via adaptive normalization of the SSVEP. Training of the classifier requires about 3 min. We have tested our system in real-time operation in three human subjects. Across subjects and sessions, control accuracy ranged from 80% to 100% correct with lags of 1-5 s for movement initiation and turning. We have also developed a realistic demonstration of our system for control of a moving map display (http://ti.arc.nasa.gov/).

Algorithms↗

The TREC 2004 genomics track categorization task: classifying full text biomedical documents.

BACKGROUND: The TREC 2004 Genomics Track focused on applying information retrieval and text mining techniques to improve the use of genomic information in biomedicine. The Genomics Track consisted of two main tasks, ad hoc retrieval and document categorization. In this paper, we describe the categorization task, which focused on the classification of full-text documents, simulating the task of curators of the Mouse Genome Informatics (MGI) system and consisting of three subtasks. One subtask of the categorization task required the triage of articles likely to have experimental evidence warranting the assignment of GO terms, while the other two subtasks were concerned with the assignment of the three top-level GO categories to each paper containing evidence for these categories. RESULTS: The track had 33 participating groups. The mean and maximum utility measure for the triage subtask was 0.3303, with a top score of 0.6512. No system was able to substantially improve results over simply using the MeSH term Mice. Analysis of significant feature overlap between the training and test sets was found to be less than expected. Sample coverage of GO terms assigned to papers in the collection was very sparse. Determining papers containing GO term evidence will likely need to be treated as separate tasks for each concept represented in GO, and therefore require much denser sampling than was available in the data sets. The annotation subtask had a mean F-measure of 0.3824, with a top score of 0.5611. The mean F-measure for the annotation plus evidence codes subtask was 0.3676, with a top score of 0.4224. Gene name recognition was found to be of benefit for this task. CONCLUSION: Automated classification of documents for GO annotation is a challenging task, as was the automated extraction of GO code hierarchies and evidence codes. However, automating these tasks would provide substantial benefit to biomedical curation, and therefore work in this area must continue. Additional experience will allow comparison and further analysis about which algorithmic features are most useful in biomedical document classification, and better understanding of the task characteristics that make automated classification feasible and useful for biomedical document curation. The TREC Genomics Track will be continuing in 2005 focusing on a wider range of triage tasks and improving results from 2004.

Journal Article↗

Streptomyces malaysiensis sp. nov., a new streptomycete species with rugose, ornamented spores.

The taxonomic position of a streptomycete strain isolated from Malaysian soil was established using a polyphasic approach. The organism, designated strain ATB-11T, was found to have chemical and morphological properties consistent with its classification in the genus Streptomyces. An almost complete 16S rRNA gene (rDNA) sequence determined for the test strain was compared with those of previously studied streptomycetes by using two treeing algorithms. The 16S rDNA sequence data not only supported classification of the strain in the genus Streptomyces but also showed that it formed a distinct phyletic line. At maturity, the aerial hyphae of strain ATB-11T differentiated into tight, spiral chains of rugose, cylindrical spores. The organism was readily distinguished from representatives of validly described Streptomyces species with rugose spores by using a combination of phenotypic features. It is proposed, therefore, that strain ATB-11T be classified in the genus Streptomyces as Streptomyces malaysiensis sp. nov.

Bacterial Typing Techniques↗

Sequence representation and prediction of protein secondary structure for structural motifs in twilight zone proteins.

Characterizing and classifying regularities in protein structure is an important element in uncovering the mechanisms that regulate protein structure, function and evolution. Recent research concentrates on analysis of structural motifs that can be used to describe larger, fold-sized structures based on homologous primary sequences. At the same time, accuracy of secondary protein structure prediction based on multiple sequence alignment drops significantly when low homology (twilight zone) sequences are considered. To this end, this paper addresses a problem of providing an alternative sequences representation that would improve ability to distinguish secondary structures for the twilight zone sequences without using alignment. We consider a novel classification problem, in which, structural motifs, referred to as structural fragments (SFs) are defined as uniform strand, helix and coil fragments. Classification of SFs allows to design novel sequence representations, and to investigate which other factors and prediction algorithms may result in the improved discrimination. Comprehensive experimental results show that statistically significant improvement in classification accuracy can be achieved by: (1) improving sequence representations, and (2) removing possible noise on the terminal residues in the SFs. Combining these two approaches reduces the error rate on average by 15% when compared to classification using standard representation and noisy information on the terminal residues, bringing the classification accuracy to over 70%. Finally, we show that certain prediction algorithms, such as neural networks and boosted decision trees, are superior to other algorithms.

Algorithms↗

A new hierarchical classification of causes of infant deaths in England and Wales.

In 1986 The Office of Population Censuses and Surveys (OPCS) introduced new certificates for stillbirths and neonatal deaths. This allowed certifiers more flexibility in the completion of the certificate, and the number and ordering of the causes given. Tabulations have been published of the fetal and maternal causes of death mentioned on the certificates for every year from 1986 to 1991 in annual reference volumes. It has not been possible either to derive a single cause group for each death, however, or to compare the information available on neonatal deaths with that on postneonatal deaths, which are still derived from the standard death certificate. The aim of the work described here was to adapt previous classifications to derive a single cause grouping for stillbirths and infant deaths which would provide the maximum information about preventability and yet meet the national and international responsibilities of OPCS. The methods used and the tests carried out on the validity and consistency of the chosen classification are described.

Algorithms↗

Analysis of tear protein patterns by a neural network as a diagnostical tool for the detection of dry eyes.

The electrophoretic patterns of tears from patients with dry-eye disease (n = 43) and from healthy subjects (n = 17) were analyzed by means of multivariate statistical methods and an artificial neural network (ANN), following sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE). From each electrophoretic pattern a data set was created, randomly divided into test (unknown samples) and training patterns (known samples), with ANN training by one of these sets. After training, the performance of the ANN was checked by presenting the test data set to the ANN. Furthermore, the data was classified using multivariate analysis of discriminance. The groups were significantly different from each other (P<0.05). The statistical procedure yielded 97% (known samples) and 71% (unknown samples) correct classifications. The ANN revealed 89% of correct classifications using the test set (unknown samples). The use of pruning algorithms (optimization procedure which automatically eliminates small weighted neurons) or genetic algorithms (optimization procedure which performs genetically induced changes of the neural net) resulted in a slight decrease of correct classifications compared to those of the nonoptimized neural network. The results reveal significant differences between the two groups. Using the ANN we were able to classify the electrophoretic tear protein pattern for diagnostic purposes.

Dry Eye Syndromes↗

Is neural network better than statistical methods in diagnosis of acute appendicitis?

Three statistical classification methods: discriminant analysis, logistic regression analysis and cluster analysis were compared with the back-propagation neural network algorithm in the diagnosis of acute appendicitis. The differences in the classification accuracy, which were evaluated with the receiver operating characteristic (ROC) curve were small, though discriminant analysis and back-propagation showed slightly better results than the other methods. The agreement of the methods on the diagnosis increased the accuracy of the classification, so that the number of misclassified cases reduced. The back-propagation neural network offers a good choice for statistical classification methods, but it was not found to be better than them. The use of several methods and their agreement as the basis of the diagnosis seems to give the best results for this diagnostic classification problem.

Appendicitis↗