PubMed HealthSearch

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Automated severity classification of AIDS hospitalizations.

To validate an automated AIDS severity-of-illness prognostic algorithm, 2,113 discharge summaries of HIV-infected patients were merged with the Problem-Oriented Medical Synopsis (POMS) and an HIV risk registry. The combination of a medically derived classification and staging algorithm with multivariate statistical techniques was used for automated severity-of-illness disease staging and prognostic assignment. The model correctly predicted the outcomes of 82% of all cases (death, survivorship) at discharge, and 66% of deaths.

Algorithms

Fine-grained structural classification of biosynthetic gene cluster-encoded products.

MOTIVATION: Biosynthetic gene clusters (BGCs) are responsible the biosynthesis of many natural products, including a multitude of effective therapeutics and their precursors. Advances in genomic data collection as well as computational techniques have made it possible to identify BGCs at scale. However, accurately determining the types of BGC-encoded products from genomic content remains elusive. RESULTS: Here, we introduce BGC annotation tool (BGCat), a machine learning method for fine-grained structural classification of BGC-encoded products, leveraging the NPClassifier natural product nomenclature. Our method leverages a pre-trained protein language model for creating meaningful gene representations and a deep neural network for class label prediction. We show the method outperforms state-of-the-art approaches in coarse-grained product classification and is effective for detailed classification. We implement a clustering-based augmentation strategy for BGC-product relationships, addressing a crucial gap in the available datasets. We then introduce the concept of product class profiles of gene cluster families (GCFs), associating each GCF with a probabilistic distribution of product types and offering a new perspective on GCF functions. Lastly, we use BGCat to provide new product class labels for over 100k BGCs in antiSMASH DB that presently have minimal information about their products. AVAILABILITY AND IMPLEMENTATION: The source code and trained model weights are freely available at https://github.com/HassounLab/BGCat.

Multigene Family

Phylogenetic classification of human papillomaviruses: correlation with clinical manifestations.

Human papillomaviruses (HPVs) are a heterogeneous group of small dsDNA viruses which cause a variety of proliferative epithelial lesions at specific anatomical sites. Although more than 65 different virus types have been cloned and characterized, no uniform classification system exists. In order to classify HPV DNA types, phylogenetic trees were constructed based on nucleotide sequence alignments using parsimony and distance matrix algorithms. The resulting phylogenetic trees provide a classification of the HPVs into specific groups encompassing the known tissue tropism and oncogenic potential of each HPV type. The implications of a phylogenetic taxonomy on the diagnostic detection of HPVs and the concept of different HPV species are discussed.

Algorithms

[The proliferative activity of myelokaryocytes and the cellular composition of the bone marrow].

Flow cytometry was used to study myelokaryocyte distribution according to the stage of the cellular cycle in 167 bone marrow specimens 94 of which were taken by puncture from hemoblastosis and anemia patients. The results obtained were compared with myelogram data. It has been established that the method provides stable and reliable values, irrespective of cellular composition of the puncture specimens. Basing on the recurrent algorithm of J. H. Fridman's classification, a computer program has been derived that permitted differential diagnosis to be made based on the data of cytometry and myelogram.

Algorithms

Searching protein sequence libraries: comparison of the sensitivity and selectivity of the Smith-Waterman and FASTA algorithms.

The sensitivity and selectivity of the FASTA and the Smith-Waterman protein sequence comparison algorithms were evaluated using the superfamily classification provided in the National Biomedical Research Foundation/Protein Identification Resource (PIR) protein sequence database. Sequences from each of the 34 superfamilies in the PIR database with 20 or more members were compared against the protein sequence database. The similarity scores of the related and unrelated sequences were determined using either the FASTA program or the Smith-Waterman local similarity algorithm. These two sets of similarity scores were used to evaluate the ability of the two comparison algorithms to identify distantly related protein sequences. The FASTA program using the ktup = 2 sensitivity setting performed as well as the Smith-Waterman algorithm for 19 of the 34 superfamilies. Increasing the sensitivity by setting ktup = 1 allowed FASTA to perform as well as Smith-Waterman on an additional 7 superfamilies. The rigorous Smith-Waterman method performed better than FASTA with ktup = 1 on 8 superfamilies, including the globins, immunoglobulin variable regions, calmodulins, and plastocyanins. Several strategies for improving the sensitivity of FASTA were examined. The greatest improvement in sensitivity was achieved by optimizing a band around the best initial region found for every library sequence. For every superfamily except the globins and immunoglobulin variable regions, this strategy was as sensitive as a full Smith-Waterman. For some sequences, additional sensitivity was achieved by including conserved but nonidentical residues in the lookup table used to identify the initial region.

Algorithms

Difficulties in diagnosing hypertension: implications and alternatives.

OBJECTIVE: To estimate the magnitude of misclassification rates with commonly used algorithms for the detection of hypertensives and to suggest a sequential approach to screening. DESIGN: A conventional statistical model was used with several different algorithms to determine the number and types of errors made in categorizing two different populations, a general population sample and a population with a high risk of hypertension. METHODS: The calculations were made for single-visit screens, similar to those used in epidemiologic studies, for three-visit screens commonly used in clinical practice and clinical trials for cutoff points of 85, 95 and 105 mmHg. A sequential probability ratio screen was proposed and the error rates estimated. RESULTS: Perhaps only one-third to two-thirds of people whose measured diastolic pressures exceed 95 mmHg actually have average pressures that high. The disparity between a single measured diastolic pressure and the mean of many pressure values also leads to errors in identifying individual subjects with mild hypertension. In a general population, single measurements of diastolic pressure exceed 95 mmHg in approximately equal numbers of normotensive, borderline and hypertensive subjects; moreover, one-third of those who are usually in the hypertensive range are not identified. All commonly used screening algorithms give too many false-positive and/or false-negative results. A sequential screening algorithm averaged 3.8 visits per subject and identified 95% of the hypertensives, with only 2.5% of those identified having usual diastolic pressures below 90 mmHg. CONCLUSIONS: Population-based surveys like the National Health and Nutrition Examination Survey (NHANES) may markedly overestimate the true prevalence of hypertension. This overestimate is greatest for mild hypertension and could significantly affect the cost/benefit analyses of public health policy. Alternative screening methods, such as the sequential algorithm proposed, may have significant benefits in providing a correct classification.

Adult

Probabilistic inference-based classification applied to myoelectric signal decomposition.

A new probabilistic inference based technique (IBC) for the classification of motor unit action potentials (MUAP's) is presented. This new technique discovers statistically significant relationships in the data and uses these relationships to generate classification rules. The technique was applied to the classification of MUAP's extracted from simulated myoelectric signals. Its performance was compared to that of classical template matching algorithms (TBC) applied to the same data. Using 32 time samples as features to represent the MUAP's it was found that the IBC based technique performed significantly better (p less than 0.005) than the TBC algorithms (83.0 +/- 2.6% versus 78.1 +/- 2.8% peak correct classification performance). As the size of the training set was reduced or as increasing numbers of random classification errors were introduced into the training data, the performance of the IBC and TBC techniques declined similarly. IBC performance remained superior until very small training sets (less than 30 MUAP's per motor unit) or training sets with large numbers of errors (greater than 50%) were used. Because the probabilistic inference technique can utilize nominal data it has the potential to use declarative problem domain knowledge which conceivably could improve its performance.

Action Potentials

Image segmentation in digital mammography: comparison of local thresholding and region growing algorithms.

Local thresholding and region-growing algorithms are developed and applied to digitized mammograms to quantify the parenchymal densities. The algorithms are first evaluated and optimized on phantom images reflecting varying image contrast, X-ray exposure conditions, and time-related changes. The difference between the segmentation results of the two techniques is less than 6% on the phantom images and 11% on the mammograms. The agreement between the computerized procedures and a manual one is in the range of 74-98%, depending on the breast parenchymal pattern and segmentation algorithm. The results show that computerized parenchymal classification of digitized mammograms is possible and independent of exposure.

Algorithms

Automatic segmentation and classification of ionic-channel signals.

Identification of ionic-channel types and their selectivity depends critically on the open channel current that can be resolved. In this paper, an automatic channel detection algorithm is proposed that is based on sequential minimization of an index which is usually used in cluster analysis. The algorithm consists of two stages, namely segmentation and classification. In the first stage, the signal samples are segmented based on the assumption that the samples in each segment should be sequentially connected. In the second stage, the resultant segments are classified with no regard to their connectivities. Results on synthetic and real channel currents are very encouraging and they suggest that this algorithm will substantially increase the productivity of many laboratories involved in ionic-channel research.

Algorithms

Image processing system for interpreting motion in American Sign Language.

In this paper, an image processing algorithm is presented for the interpretation of the American Sign Language (ASL), which is one of the sign languages used by the majority of the deaf community. The process involves detection of hand motion, tracking the hand location based on the motion and classification of signs using adaptive clustering of stop positions, simple shape of the trajectory, and matching of the hand shape at the stop position.

Algorithms

[Rough sets theory in the analysis of structure-activity relationships of quaternary quinolinium- and isoquinolinium compounds].

Relationship between chemical structure and antimicrobial activity of 72 quaternary quinolinium and isoquinolinium compounds is analyzed using the theory of rough sets. The compounds are described by 11 attributes concerning structure and are divided into 3 classes of activity. The description builds up on information system. Using the rough sets approach a smallest set of attributes significant for a high quality of classification has been found. A decision algorithm has been driven from the information system showing up important relations between structure and activity. This may be helpful in supporting decisions concerning synthesis of new antimicrobial compounds.

4-Quinolones

[Rough sets theory in structure-activity relationship analysis of quaternary pyridinium compounds].

Relationship between chemical structure and antimicrobial activity of 53 quaternary pyridinium compounds is analysed using the theory of rough sets. The compounds are described by 8 attributes concerning structure and are divided into 5 classes of activity. The description builds up an information system. Using the rough sets approach a smallest set of attributes significant for a high quality of classification has been found. A decision algorithm has been derived from the information system showing important relations between structure and activity. It may be helpful in supporting decisions concerning synthesis of new antimicrobial compounds.

Algorithms

[Semi-automatic TNM classification of malignant tumors with the ESTER system exemplified by the larynx].

Classification of tumours according to the TNM scheme has been accepted worldwide. However, vague baseline assessments and borderline cases render a comparison of the outcome on the basis of TNM classification impossible. Therefore we integrated the TNM rules as a new algorithm into an existing expert system for determining therapy. Thus, every tumour documented with the ESTHER system is automatically classified according to current TNM rules. The program is designed to cope with future changes of the TNM system: raw data are used for classification so that only the algorithms need to be modified.

Expert Systems

A national study of medical and surgical specialties. III. An empirical approach to the classification of patient care.

A major feature of a national survey of medical and surgical specialties is the development and application of an algorithm for classifying patient care services provided by physicians. The care classification reflects much of prevailing opinion regarding what constitutes primary and nonprimary care. The classification system provides a powerful tool for the analysis of patient care services, since it is based on conditions of access to care, the physician's role in providing the care, measures associated with continuity of care, and a proxy measure of comprehensiveness of care. Furthermore, it is based on the recordings by physicians of actual patient-encounter characteristics and is not operationally dependent on physician characteristics or propensities.

Cardiology

Hierarchical Multi-Label Classification With Gene-Environment Interactions in Disease Modeling.

In biomedical studies, gene-environment (G-E) interactions have been demonstrated to have important implications for analyzing disease outcomes beyond the main G and main E effects. Many approaches have been developed for G-E interaction analysis, yielding important findings. However, hierarchical multi-label classification, which provides insightful information on disease outcomes, remains unexplored in G-E analysis literature. Moreover, unlabeled data are commonly observed in practical settings but omitted by many existing methods of hierarchical multi-label classification. In this study, we consider a semi-supervised scenario and develop a novel approach for the two-layer hierarchical response with G-E interactions. A two-step penalized estimation is then proposed using an efficient expectation-maximization (EM) algorithm. Simulation shows that it has superior performance in classification and feature selection. The analysis of The Cancer Genome Atlas (TCGA) data on lung cancer demonstrates the practical utility of the proposed method. Overall, this study can fill the important knowledge gap in G-E interaction analysis by providing a widely applicable framework for hierarchical multi-label classification of complex disease outcomes.

Humans

[Numerical taxonomy of the genus Desulfovibrio by group analysis].

The Desulfovibrio genus has a particular interest because it includes the microorganisms connected with the corrosion produced microbiologically. The taxonomy of the genus shows disadvantages due to its metabolical and physiological characteristics. In this paper, 14 strains of the Desulfovibrio type were studied from the metabolical point of view. Numeric taxonomy was carried out according to the Group Analysis method, using and comparing the change possibilities of the method. The Consensus Method was also applied. The results obtained indicate a low metabolic activity of the strains with regard to the number of compounds which can be used as energy source. The taxonomic method showed a better structure with more clear divisions, corresponding to Simple Matching coefficient (which coincides with other symmetric coefficients and with the distance coefficient) with average bond (UPGMA). It is estimated that the present classification will vary in time with new strains with different metabolic characteristics. The two groups of bacteria correspond to those with more and less degrading ability.

Algorithms

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans

Structure analysis and classification of cervical cells using a processing system based on TV.

This paper presents preliminary results of a cell classification experiment using a new approach for feature extraction. The algorithm takes into account the special requirements of a fast parallel processing system (processor-oriented algorithms). A cell image is described by several hundred features derived from the nucleus only. The most significant features with respect to classification are determined by statistical analysis. Applying principal axis transform, a new feature set is computed, reduced considerably in dimensions. The data base (1,925 cell images of Papanicolaou-stained cervical specimens) was divided into a training set (963 images) and a test set (962 images). The classification results of the test set show that the recognition rate for the two-class problem (normal, suspicious) is better than 91%, using only ten morphologic features.

Cervix Mucus