PubMed HealthSearch

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Hierarchical Multi-Label Classification With Gene-Environment Interactions in Disease Modeling.

In biomedical studies, gene-environment (G-E) interactions have been demonstrated to have important implications for analyzing disease outcomes beyond the main G and main E effects. Many approaches have been developed for G-E interaction analysis, yielding important findings. However, hierarchical multi-label classification, which provides insightful information on disease outcomes, remains unexplored in G-E analysis literature. Moreover, unlabeled data are commonly observed in practical settings but omitted by many existing methods of hierarchical multi-label classification. In this study, we consider a semi-supervised scenario and develop a novel approach for the two-layer hierarchical response with G-E interactions. A two-step penalized estimation is then proposed using an efficient expectation-maximization (EM) algorithm. Simulation shows that it has superior performance in classification and feature selection. The analysis of The Cancer Genome Atlas (TCGA) data on lung cancer demonstrates the practical utility of the proposed method. Overall, this study can fill the important knowledge gap in G-E interaction analysis by providing a widely applicable framework for hierarchical multi-label classification of complex disease outcomes.

Humans

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans

Structure analysis and classification of cervical cells using a processing system based on TV.

This paper presents preliminary results of a cell classification experiment using a new approach for feature extraction. The algorithm takes into account the special requirements of a fast parallel processing system (processor-oriented algorithms). A cell image is described by several hundred features derived from the nucleus only. The most significant features with respect to classification are determined by statistical analysis. Applying principal axis transform, a new feature set is computed, reduced considerably in dimensions. The data base (1,925 cell images of Papanicolaou-stained cervical specimens) was divided into a training set (963 images) and a test set (962 images). The classification results of the test set show that the recognition rate for the two-class problem (normal, suspicious) is better than 91%, using only ten morphologic features.

Cervix Mucus

acmgscaler: an R package and Colab for standardized gene-level variant effect score calibration within the ACMG/AMP framework.

MOTIVATION: A genome-wide variant effect calibration method was recently developed under the guidelines of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP), following ClinGen recommendations for variant classification. While genome-wide approaches offer clinical utility, emerging evidence highlights the need for gene- and context-specific calibration to improve accuracy. Building on previous work, we have developed an algorithm tailored to converting functional scores from both multiplexed assays of variant effects (MAVEs) and computational variant effect predictors (VEPs) into ACMG/AMP evidence strengths. RESULTS: Our method is designed to deliver consistent performance across different genes and score distributions, with all variables adaptively determined from the input data, preventing selective adjustments or overfitting that could inflate evidence strengths beyond empirical support. To facilitate adoption, we introduce acmgscaler, a lightweight R package and a plug-and-play Google Colab notebook for the calibration of custom datasets. This algorithmic framework bridges the gap between MAVEs/VEPs and clinically actionable variant classification. AVAILABILITY AND IMPLEMENTATION: The R package and Colab notebook are available at https://github.com/badonyi/acmgscaler.

Software

[Interpretation of pulse curves by means of the Walsh analysis].

A method for classification of medical data using the Walsh transformation is demonstrated. The Walsh spectrum was obtained by the algorithm of Andrews-Kane-Pratt. The spectral points were used to declare the signals normal or abnormal. The example described in this paper shows that in the case of pulse wave the classification is successful in 93 percent.

Carotid Arteries

engGNN: a dual-graph neural network for omics-based disease classification and feature selection.

Omics data, such as transcriptomics, proteomics, and metabolomics, provide critical insights into disease mechanisms and clinical outcomes. However, their high dimensionality, small sample sizes, and intricate biological networks pose major challenges for reliable prediction and meaningful interpretation. Graph neural networks offer a promising way to integrate prior knowledge by encoding feature relationships as graphs. Yet, existing methods typically rely solely on either an externally curated feature graph or a data-driven generated graph, which limits their ability to capture complementary information. To address this, we propose the external and generated Graph Neural Network (engGNN), a dual-graph framework that jointly leverages both external biological networks and data-driven generated graphs. Specifically, engGNN constructs a biologically informed undirected feature graph from established network databases and complements it with a directed feature graph derived from tree-ensemble models. This dual-graph design produces more comprehensive representations, thereby improving predictive performance and interpretability. Through extensive simulation studies and real-world applications to three independent gene expression datasets, engGNN consistently demonstrates strong classification performance compared with competitive baselines. Beyond classification, engGNN provides feature- and source-level interpretability, enabling biologically meaningful analyses such as pathway enrichment analysis. Taken together, these results highlight engGNN as a robust, flexible, and interpretable framework for disease classification and biomarker discovery in high-dimensional omics contexts.

Graph Neural Networks

Identifying JAK2 and ANXA5 as Key Genes Linking Obstructive Sleep Apnea and Oxidative Stress via Machine Learning and Multilayer Transcriptomic Integration With Functional Validation.

Obstructive sleep apnea (OSA) is a common and severe sleep disorder closely associated with oxidative stress (OS). This study aims to identify and validate potential OS-related genes associated with OSA through bioinformatics methods. We successfully identified OS-related differentially expressed genes (OS-DEGs) by combining the limma test, weighted correlation network analysis (WGCNA), and OS-related genes from the GeneCards database. Key genes and potential biological roles were further identified using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), enrichment analysis, protein-protein interaction (PPI) network analysis, Lasso regression analysis, random forest algorithm, and support vector machine recursive feature elimination (SVM-RFE) method. Evaluate and validate the accuracy of key genes through receiver operating characteristic (ROC) curve analysis. The human single-cell RNA sequencing (scRNA-seq) dataset is used for cell classification annotation, analysis of key gene single-cell expression profiles, and virtual gene knockout experiments based on the scTenifoldKnk algorithm. Integrating scRNA-seq sequencing, pseudotime trajectory inference, cell-cell communication analysis, and bulk immune infiltration deconvolution reveals monocyte subtype remodeling in OSA. Finally, the expression levels of key genes in clinical samples were validated using real-time quantitative PCR (RT-qPCR) and Western blotting. A total of 57 common DEGs, indicating significant enrichment in OS, inflammation, and tumor pathways, particularly prominent in the immunometabolism pathway. By integrating DEGs, WGCNA, PPI results, and machine learning methods, key genes Janus kinase 2 (JAK2) and ANXA5 were screened out. JAK2 was significantly upregulated under disease conditions, while ANXA5 was significantly downregulated. ROC curve exhibited high accuracy (area under the curve [AUC] > 0.85). Human scRNA-seq analysis revealed that key genes were predominantly highly expressed in monocytes. Virtual knockout experiments demonstrated that these key genes play a crucial role in regulating immune responses and inflammatory reactions. PPI networks and enrichment analysis verified that downstream genes S100P, ALOX5AP, PROK2, and PADI4 may collaboratively participate in immune response and inflammation regulation. Finally, clinical sample experiment further validated the results of bioinformatics analysis. This study provides new research insights for the diagnosis, mechanism research, and treatment development of OSA in the future by integrating multilayer transcriptomic and machine learning techniques.

Humans

VISTA: a classifier for metagenomic subspecies and community state typing of the vaginal microbiome.

Metagenomic community state types (mgCSTs) capture within-species genetic and functional diversity and community structure of the vaginal microbiome, enabling precise links between microbiome composition, function, and health-related risk. VISTA, the Vaginal Inference of Subspecies and Typing Algorithm, is a two-step classifier that assigns mgCSTs to vaginal metagenomes, providing standardized, scalable classifications.

bioinformatics

Discovery and performance of DNA methylation panels for cancer detection and classification in blood.

Examining DNA in a liquid biopsy for non-invasive cancer detection relies on identifying dilute signal in a high background. This study aims to identify DNA methylation biomarkers for multi-cancer detection. Utilizing large tissue datasets, we apply novel search algorithms to discover confined biomarker panels capable of distinguishing tumor from normal and determining the tissue of origin. We explore the applicability to blood-based testing using targeted methylation sequencing followed by machine learning classification. We present an 8-marker panel, which successfully predicts tumors across 14 types with a 91% average sensitivity, maintaining a low false positive rate (< 0.04%). Additionally, a panel of 39 CpG sites exhibits accuracies ranging from 69% to 98% for identifying tissue of origin. When tested on 114 patient plasma samples (colon, liver, pancreatic, prostate, and stomach cancer), the 8-marker panel obtains an AUC of 0.78 with a 78% sensitivity among 32 early-stage patients (stage I-II), and 60% overall. Using the 39-marker panel in a multi-class classification model selecting only the best match, 54% of tumor samples were on average correctly assigned to the tissue of origin, and up to 80% when allowing more inclusive criteria. Using a limited set of biomarkers, our work contributes to advancing non-invasive cancer diagnostics.

DNA methylation

MegaPX: fast and space-efficient peptide assignment method using IBF-based multi-indexing.

MOTIVATION: A central problem for metaproteomic analysis is the often-unknown taxonomic composition of the analyzed microbiomes. Using a database search, the standard approach requires prior knowledge of which proteins and taxa to include in the protein reference database or to use tailored metagenome-derived databases, which are expensive and error-prone in their generation. A possible strategy to circumvent this database search issue is de novo sequencing, where peptide sequences are directly identified from mass spectra. However, these sequences must still be mapped back to potentially extensive databases. Here, alignment-based approaches enable robust and precise results, with the potential drawback of high memory usage and long run times. RESULTS: We present MegaPX, a software for rapidly classifying de novo peptide sequences against large protein databases. MegaPX implemented as a C++-based tool, uses an alignment-free, k-mer approach as a taxonomic classification method with the possibility of generating mutated reference databases for error-tolerant searching. It uses various algorithms, including interleaved Bloom filters, to efficiently compute approximate membership queries, ensuring fast processing times while querying and indexing large databases in a multi-indexing fashion. We demonstrate the potential of MegaPX by analyzing different samples, including metaproteomics, against extensive reference databases, highlighting its use as a fast screening tool.

Software

Echocardiogram analysis in a pattern recognition framework.

Echocardiogram analysis is treated in a pattern recognition framework. Anterior mitral leaflet waveforms are classified for the four-class problem consisting of the classes "normal," "mitral stenosis," "mitral valve prolapse," and "idiopathic hypertrophic subaortic stenosis." In addition, aortic root waveforms and left ventricular wall waveforms are classified for the two-class problem consisting of the classes "normal" and "idiopathic hypertrophic subaortic stenosis." One common method of analysis (Fourier analysis) underlies each classification scheme. Classification accuracy is sufficiently good to warrant the inference that successful automated decision-making based on the algorithms investigated is feasible.

Aortic Valve

Automated Classification of Lymphoma Subtypes From Histopathological Images Using a U-Net Deep Learning Model: Comparative Evaluation Study.

BACKGROUND: Accurate classification and grading of lymphoma subtypes are essential for treatment planning. Traditional diagnostic methods face challenges of subjectivity and inefficiency, highlighting the need for automated solutions based on deep learning techniques. OBJECTIVE: This study aimed to investigate the application of deep learning technology, specifically the U-Net model, in classifying and grading lymphoma subtypes to enhance diagnostic precision and efficiency. METHODS: In this study, the U-Net model was used as the primary tool for image segmentation integrated with attention mechanisms and residual networks for feature extraction and classification. A total of 620 high-quality histopathological images representing 3 major lymphoma subtypes were collected from The Cancer Genome Atlas and the Cancer Imaging Archive. All images underwent standardized preprocessing, including Gaussian filtering for noise reduction, histogram equalization, and normalization. Data augmentation techniques such as rotation, flipping, and scaling were applied to improve the model's generalization capability. The dataset was divided into training (70%), validation (15%), and test (15%) subsets. Five-fold cross-validation was used to assess model robustness. Performance was benchmarked against mainstream convolutional neural network architectures, including fully convolutional network, SegNet, and DeepLabv3+. RESULTS: The U-Net model achieved high segmentation accuracy, effectively delineating lesion regions and improving the quality of input for classification and grading. The incorporation of attention mechanisms further improved the model's ability to extract key features, whereas the residual structure of the residual network enhanced classification accuracy for complex images. In the test set (N=1250), the proposed fusion model achieved an accuracy of 92% (1150/1250), a sensitivity of 91.04% (1138/1250), a specificity of 89.04% (1113/1250), and an F1-score of 90% (1125/1250) for the classification of the 3 lymphoma subtypes, with an area under the receiver operating characteristic curve of 0.95 (95% CI 0.93-0.97). The high sensitivity and specificity of the model indicate strong clinical applicability, particularly as an assistive diagnostic tool. CONCLUSIONS: Deep learning techniques based on the U-Net architecture offer considerable advantages in the automated classification and grading of lymphoma subtypes. The proposed model significantly improved diagnostic accuracy and accelerated pathological evaluation, providing efficient and precise support for clinical decision-making. Future work may focus on enhancing model robustness through integration with advanced algorithms and validating performance across multicenter clinical datasets. The model also holds promise for deployment in digital pathology platforms and artificial intelligence-assisted diagnostic workflows, improving screening efficiency and promoting consistency in pathological classification.

Humans

GeomeTRe: accurate calculation of geometrical descriptors of tandem repeat proteins.

MOTIVATION: Structured tandem repeat proteins (STRPs) are characterized by preserved structural motifs arranged in a modular way. The structural and functional diversity of STRPs makes them particularly important for studying evolution and novel structure-function relationships, and ultimately for designing new synthetic proteins with specific functions. One crucial aspect of their classification is the estimation of geometrical parameters, which can provide better insight into their properties and the relationship between the spatial arrangement of repeated units and protein function. Calculating geometric descriptors for STRPs is challenging because naturally occurring repeats are not "perfect" and often contain insertions and deletions. Existing tools for predicting structural symmetry work well on simple cases but often fail for most natural proteins. RESULTS: Here, we present GeomeTRe, an algorithm that calculates geometrical descriptors such as curvature (yaw), twist (roll), and pitch for a protein structure with known repeat unit positions. The algorithm simulates the movement of consecutive units, identifies rotational axes, and calculates the corresponding Tait-Bryan angles. GeomeTRe's parameters can enhance STRP annotation and classification by identifying variations in geometric arrangements among different functional groups. The package is fast and suitable for processing large protein structure datasets when repeat region information (e.g. from RepeatsDB) is available. AVAILABILITY AND IMPLEMENTATION: GeomeTRe is available as a Python package; source code and documentation can be found at https://github.com/BioComputingUP/GeomeTRe.

Algorithms

Computer interpreted fetal electroencephalogram: sharp wave detection and classification of infants for one year neurological outcome.

The presence of visually discernible sharp waves (SWs) in the fetal electroencephalogram (FEEG) has been found to be associated with abnormal neurological infant outcome, but no method of programmed SW detection for FEEG was available. In order to develop an algorithm for SW detection, the first and second derivatives for visually identified SWs and non-SWs were examined and five random variables chosen for discriminant function analysis (DFA). The resulting equation, incorporated into program logic along with logic for artifact rejection, produced classifications from 85% to 89% consistent with visual identifications, suggesting that the number of SWs/epoch (NSW) corresponds with visually identified SWs. In addition, in 61 cases using a threshold for NSW derived by DFA, computer recognized SWs were found to be significantly related to the overall visual interpretation of the tracings (P less than 0.005). Finally, NSW alone produced correct classification of 65.5% of infants for 1 year neurological outcome. The overall consistency was increased to as high as 80% using additional FEEG and neonatal data. These findings imply that some forms of brain damage are present before birth and can be detected during labor using FEEG.

Brain Diseases

Digital and computational morphology in hematology: current platforms, clinical evidence, and future requirements.

INTRODUCTION: Morphologic examination of peripheral blood and bone marrow remains central to the diagnosis and classification of hematologic disorders. Conventional optical microscopy, however, is labor-intensive, dependent on operator expertise, and affected by interobserver variability. Digital morphology has developed from automated image acquisition and cell pre-classification into a broader field that includes whole-slide imaging, remote review, quantitative morphometry, and artificial intelligence-based analysis. CONTENT: This review examines current applications of digital morphology in peripheral blood, bone marrow aspirates, malaria detection, and body-fluid analysis. Commercial platforms are evaluated with particular attention to the distinction between raw automated pre-classification, expert digital post-classification, and comparison with independent optical microscopy. Digital systems generally perform well for common mature leukocyte populations but remain less reliable for rare or diagnostically critical cells, including blasts, abnormal lymphoid cells, plasma cells, and intermediate maturation stages. Research systems increasingly extend analysis from individual-cell classification to whole-slide, specimen-level, and patient-level assessment. SUMMARY: Digital morphology can improve standardization, image traceability, remote consultation, education, proficiency testing, quality assurance, and selected aspects of laboratory workflow. Its clinical value depends on appropriate validation, transparent reporting of reference methods, recognition of algorithm-specific failure modes, and clearly defined criteria for expert review and conventional microscopy. Human expertise remains essential not only for validating results but also for adapting cell taxonomies and interpretive rules to evolving classifications of hematologic diseases. OUTLOOK: Future progress will require representative multicenter datasets, harmonized morphologic terminology, external validation, interoperability with laboratory information systems, and continuous monitoring after software or hardware updates. Integration of morphology with quantitative hematology, flow cytometry, cytogenetics, genomics, and clinical data may support more comprehensive computational diagnosis. Digital platforms may also broaden access to specialist expertise, training, and quality programs in resource-limited institutions and regions, provided that infrastructure, governance, and professional competency are adequately supported.

artificial intelligence

An Integrative Morphological and Genomic Analysis With a Refined Fluorescence In Situ Hybridization (FISH) Threshold and Novel Kinase Fusions in a Large Asian Cohort of Spitzoid Neoplasms.

Differentiating atypical Spitz tumors (ASTs) from true Spitz melanomas (SMs) and conventional melanomas with spitzoid features (MSFs) remains a formidable diagnostic challenge. Because current molecular epidemiological data are overwhelmingly derived from Caucasian cohorts, the genomic landscape of Asian populations remains largely unexplored. To elucidate the molecular progression landscape and refine the diagnostic criteria, we performed a comprehensive multimodal analysis-integrating histomorphology, immunohistochemistry, multiprobe fluorescence in situ hybridization (FISH), and targeted RNA/DNA-based next-generation sequencing (NGS)-on a cohort of 140 spitzoid neoplasms. This cohort, comprising 126 ASTs, 8 SMs, and 6 MSFs, represents the largest Asian cohort to date. Malignant phenotype strongly correlated with lesional asymmetry, deep atypical mitoses, a sheet-like growth pattern, diffuse preferentially expressed antigen of melanoma positivity, and significant loss of p16 expression (64.3% in SM/MSF vs 9.5% in ASTs; P < .0001). Building upon the established melanoma FISH criteria, we optimized a prognostic threshold of &#x2265;2 FISH abnormalities specifically tailored for spitzoid neoplasms. We demonstrated that isolated single chromosomal aberrations (particularly MYB loss) are relatively stable events that are frequent in indolent ASTs, whereas our refined &#x2265;2 threshold yielded 100% sensitivity and 92.5% specificity for predicting regional lymph node metastasis/local recurrence. Molecularly, NGS identified mutually exclusive initiating driver alterations (comprising kinase fusions and HRAS mutations) in 89.9% of true Spitz neoplasms, a remarkably high prevalence suggesting a distinct genetic background in Asian populations. We also characterized 5 entirely novel kinase fusions (ZNF24::ROS1, PCBP1::ROS1, NUMA1::RET, CBWD1::ALK, and TPR::NTRK1). Furthermore, NGS definitively segregated true Spitz neoplasms from morphological mimics (MSF), which lacked fusions and were driven by canonical genomic alterations of the conventional melanoma pathway. Integrating these genomic landscapes validated a stepwise progression model. Although isolated kinase fusions drove indolent ASTs, malignant SM invariably harbored concurrent pathogenic secondary alterations, demonstrating a profound reliance on CDKN2A/B, TP53, and CDK4 aberrations. Ultimately, we propose an integrated diagnostic algorithm combining morphological evaluation, the refined FISH threshold, and comprehensive NGS profiling, providing a precise, evidence-based framework for pathway classification and clinical management of spitzoid neoplasms.

fluorescence in situ hybridization