PubMed HealthSearch

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Single sweep analysis of visual evoked potentials through a model of parametric identification.

An original method is presented for the single sweep analysis of visual evoked potentials (VEP's). The introduced algorithm bases upon an AutoRegressive with eXogenous input (ARX) modeling. A Least Squares procedure estimates the coefficients of the model and allows to obtain a complete black-box description of the signal generation mechanism, besides providing a filtered version of the single sweep potential. The performance of the algorithm is verified on proper simulation tests and the experimental results put into evidence the noticeable improvement of signal-to-noise ratio with a consequent better recognition of the classical parameters of the peaks (latencies and amplitudes). The possibility of measuring these parameters on a single sweep basis enables to evaluate the dynamics of the Central Nervous System response during the entire course of the examination. A classification of the estimated evoked potentials in a small number of subsets, on the basis of their morphology, is also possible.

Algorithms

Acoustic classification of alarm calls by vervet monkeys (Cercopithecus aethiops) and humans (Homo sapiens): II. Synthetic calls.

In 2 experiments classification of synthetic versions of species-typical snake and eagle alarm calls by vervet monkeys (Cercopithecus aethiops) and human (Homo sapiens) control subjects was investigated. In a 2-choice, operant-conditioning-based procedure, this work followed up acoustic analyses that had used various digitally based algorithms (Owren & Bernacki, 1988). All subjects were first tested with alarm-call replicas that were based on analysis data. These models were classified in the same manner as natural stimuli, which verified the appropriateness of the acoustic characterizations. Synthetic stimuli were then presented to test the importance of specific acoustic cues. Spectral patterning was found to be the most salient cue for classification by the monkeys, whereas results from the human subjects were mixed. Implications for the study of nonhuman primate vocalizations and Lieberman's (1984) theory of speech evolution are discussed.

Adult

DeepLabCut-based automated system reveals diverse temperature tolerance among medaka strains and related Oryzias species.

Temperature is a critical environmental factor influencing the physiology and behavior of ectothermic animals, yet conventional methods for evaluating thermal tolerance in fish rely on subjective manual observation of loss of equilibrium (LOE), limiting experimental throughput and introducing observer bias. Here, we developed an automated temperature tolerance evaluation system integrating DeepLabCut-based pose estimation with custom image processing algorithms to objectively quantify the timing of LOE during thermal stress tests. Our system incorporated region partitioning and color transformation preprocessing to improve keypoint detection accuracy, followed by a classification model combining ResNet34-based frame features with keypoint coordinates to objectively determine the timing of LOE without manual observation. Validation against manual annotation showed that the automated system achieved an accuracy comparable to the natural variability between trained investigators, and outperformed naive human observers, supporting its validity as an objective and reproducible alternative to manual scoring. Using this system, we characterized cold and heat tolerance across six medaka strains (Oryzias latipes: d-rR/TOKYO, HB11A, OK-Cab, HO5 and HdrR-II1; O. sakaizumii: HNI-II). Cold and heat tolerance assessment revealed inter-strain variation, with HdrR-II1 among the most cold- and heat-tolerant strains and HNI-II the least tolerant of both cold and heat stress. We further evaluated cold tolerance in medaka-related species (O. sinensis, O. cabaranensis, O. curvinotus, O. luzonensis, O. celebensis, and O. javanicus) and zebrafish (Danio rerio), revealing substantial interspecific variation that broadly corresponded with latitudinal distribution. O. latipes, distributed at the highest latitudes among the tested species, exhibited the greatest cold tolerance, whereas O. celebensis, O. javanicus, and other tropical or low-latitude species showed comparatively low cold tolerance. Our automated system provides a robust, high-throughput platform for thermal tolerance evaluation and, combined with the genetic and genomic resources available in medaka, establishes a foundation for elucidating the molecular mechanisms underlying temperature adaptation in fish.

Animals

Treatment of lower extremity infections in diabetics.

The infected diabetic lower extremity has enjoyed a surge in popularity in the medical literature. There have been numerous papers outlining classification systems for ulcer depth, surgical approaches, and microbiology. Discussions on antibiotic use have usually been directed toward therapy of the "diabetic foot infections" as a group, without regard to differences in severity and location of these infections. These infections can vary from the most superficial of processes to a severe life- and limb-threatening sepsis. The author presents a review of the processes involved in the diabetic lower extremity infection and suggests a classification system for selection of empiric antibiotic therapy based on the severity of the infection.

Algorithms

AniAnn's: alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates.

MOTIVATION: Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. RESULTS: In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. AVAILABILITY: AniAnn's is open source software and available at github.com/marbl/anianns.

Algorithms

Comparison of the performances of an automated arrhythmia detector working on original and virtual ECG tracings.

Two approaches can be taken to improve the performance of an automatic arrhythmia detector: perfecting the detection algorithms or improving the quality of the investigated traces by preprocessing the original traces. This paper reports on the results of a data preprocessing approach. Preprocessing consists in constructing new traces, which we call virtual. They are mathematically obtained from the original traces and referred to the dominant cardiac electric axis. The classifications obtained with an arrhythmia detector using both virtual and original traces are presented and discussed. By comparing the performance indices obtained under the two different conditions, it can be seen that a diagnosis based on the virtual traces is as acceptable as one based on the original traces. This result should be judged as favorable, since the algorithm was not adjusted or calibrated to the virtual traces, while those who developed it had certainly calibrated the parameters to the original traces.

Algorithms

Tests of homogeneity and trend with medians.

Algorithms for tests of homogeneity and trend with medians are presented. Under certain conditions the tests are more efficient than many standard nonparametric tests on one-way classification designs. Computationally the tests are very simple and could be performed by hand or with the help of a calculator. Application to certain types of data derived in cytogenetics is pointed out with an example.

Animals

Cases of alleged asbestos-related disease: a radiologic re-evaluation.

Chest radiographs were re-evaluated from 439 active and retired tireworkers previously designated as having a condition consistent with an asbestiform mineral exposure. The review was performed in an independent manner by three board-certified radiologists according to guidelines from an international classification system. The percentage of cases with abnormalities consistent with an asbestiform mineral exposure found separately by the three radiologists was 3.7, 3.0, and 2.7%. Application of an algorithm to form a consensus evaluation indicated that approximately 3.6% (16) of the subjects evaluated may have a condition consistent with an asbestos exposure. A more detailed review, however, revealed that only 11 workers, or 2.5% of the total, would have a reasonable likelihood of having such a condition. Most cases were normal and the majority of abnormalities present on the radiographs evaluated were nonoccupational in origin. Prevalent conditions identified included healed tuberculosis, histoplasmosis, emphysema, discoid atelectasis, effusions, healed rib fractures, scarring due to infection or old inflammatory disease, possible cancer, miscellaneous nonspecific linear markings consistent with cigarette smoking and aging, and heart and vascular system diseases--the latter evidenced by an abnormally large number of subjects with healed coronary artery bypass surgery and pacemaker implants. In summary, the best estimate from this study indicates that possibly 16 (3.6%), but more realistically 11 (2.5%), of the 439 tireworkers evaluated may have a condition consistent with exposure to an asbestiform mineral. This represents a 40-fold difference between the re-evaluation results and the original survey work.

Asbestosis

Predicting host tropism in influenza a viruses: insights from multi-segment nucleotide signatures.

BACKGROUND: Influenza A virus (IAV) poses a significant public health threat due to its cross-species transmission and complex host adaptation mechanisms. This study integrated whole-genome data from avian, human, swine, and bovine IAV strains, using machine learning to predict viral host tropism based on nucleotide site features and to identify key sites driving host adaptation along with their synergistic effects. METHODS: A total of 64,000 IAV sequences from avian, human, swine, and bovine hosts were analyzed to build host-prediction models. A four-class classification framework (avian, human, swine, bovine) was constructed using nucleotide site features from all eight genomic segments (PB2, PB1, PA, HA, NP, NA, MP, NS). Eight machine learning algorithms (logistic regression, decision tree, random forest, SVM, KNN, gradient boosting, XGBoost, LightGBM) were benchmarked via 10-fold stratified cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, AUPRC, and AUC. SHAP (SHapley Additive exPlanations) analysis prioritized critical nucleotide sites, while bivariate association tests identified synergistic/antagonistic interactions between sites. Nucleotide composition profiles were compared across host groups using hierarchical clustering and heatmap visualization. RESULTS: The XGBoost algorithm demonstrated the best and most stable performance, achieving an AUC value of over 0.95 in distinguishing human-derived sequences from non-human ones. SHAP analysis identified the top 20 critical nucleotide sites for each gene segment, such as sites 46 and 698 in the NS segment. Nucleotide composition analysis revealed high similarity between human and swine sequences in the HA and PB2 segments, and between avian and bovine sequences. The HA segment was particularly challenging in differentiating human from swine strains. Bivariate site association analysis uncovered significant synergistic or antagonistic effects between key sites within gene segments, forming complex networks. For instance, in the NS segment, a positive prediction contribution was observed when sites 371, 698, and 419 were all G. CONCLUSIONS: This study advances our mechanistic understanding of IAV host adaptation, identifies molecular determinants for zoonotic risk stratification, and establishes a scalable machine learning framework for predicting viral host tropism through nucleotide signature analysis, thereby enhancing surveillance strategies and informing preventive measures against emerging viral threats.

Influenza A virus

Automatic wave form classification of extracellular multineuron recordings.

A PC-based method for the reconstruction of individual spike trains from extracellular multineuron recordings is described. Starting with virtually no knowledge about the wave forms in a record, a fully automatic template-finding algorithm extracts templates using the entire data set. In a second step, individual spike trains are reconstructed.

Algorithms

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans

Comparison of the ID3 algorithm versus discriminant analysis for performing feature selection.

Having obtained disappointing results in a small medical data set despite the fact that our data seemed to be well suited for induction via ID3, we decided to compare the performance of ID3 to discriminant analysis. Performance was gauged by the percentage of correct classification in a second, independent data set. Examples were obtained from a cardiology project on the accuracy of auscultation. There were 107 examples in the first data set and 67 cases in the second. We found that ID3 and discriminant analysis performed equally poorly, with ID3 classifying only 60% of the second set correctly and discriminant analysis classifying 66% of the second set correctly. Also, the ID3 probability statistic for estimating the accuracy of ID3 for classifying further cases was markedly optimistic compared to our actual second data set results. Moreover, with an increase in sample size, ID3 seemed to break down, producing a large, complex decision tree of dubious generality, whereas discriminant analysis, with a larger sample size, used more independent variables but maintained its first set accuracy. These data suggest that there is a need for more sophisticated algorithms than ID3, even at the risk of giving up some computational efficiency.

Adult

Unsupervised clustering and centroid estimation using dynamic competitive learning.

In this paper, an unsupervised learning algorithm is developed. Two versions of an artificial neural network, termed a differentiator, are described. It is shown that our algorithm is a dynamic variation of the competitive learning found in most unsupervised learning systems. These systems are frequently used for solving certain pattern recognition tasks such as pattern classification and k-means clustering. Using computer simulation, it is shown that dynamic competitive learning outperforms simple competitive learning methods in solving cluster detection and centroid estimation problems. The simulation results demonstrate that high quality clusters are detected by our method in a short training time. Either a distortion function or the minimum spanning tree method of clustering is used to verify the clustering results. By taking full advantage of all the information presented in the course of training in the differentiator, we demonstrate a powerful adaptive system capable of learning continuously changing patterns.

Cluster Analysis

Identification of the elastic symmetry of bone and other materials.

A simplified classification scheme for the elastic symmetries of a solid is applied to the identification of the elastic symmetry of a material by three different methods--visual, stereological and numerical algorithm. Each method is illustrated with an application to bone tissues, but the methods apply to all materials.

Algorithms

HCSeeker: A classification tool for human genetic variant hot and cold spots designed for PM1 and benign criteria in the ACMG-AMP guideline.

PURPOSE: The PM1 criterion, which states that a variant is located in a mutational hot spot and/or critical and well-established functional domain without benign variation (such as the active site of an enzyme), is considered moderate evidence for assessing its pathogenicity. Although guidelines from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology are widely adopted, the PM1 criterion remains limited from lacking a reliable database of variant hot spots. Compared with hot spots, cold spots are neglected by the guidelines. To improve variant classification, we suggest including cold spots for supporting benign classifications. Consequently, we have developed the HCSeeker to provide data support for PM1 and the "Benign" criteria. METHODS: HCSeeker uses the Kernel Density Estimation and the Expectation-Maximization algorithm to identify hot- and cold-spot regions. RESULTS: Through HCSeeker, we identified 988 hot spots and 682 cold spots across 889 genes and provided a public database (http://www.genemed.tech/hcseeker/) for researchers and clinicians to query variant locations, facilitating the application of American College of Medical Genetics and Genomics and the Association for Molecular Pathology PM1 or "Benign" criteria. CONCLUSION: We developed the HCSeeker tool, which can effectively identify variant hot and cold spots within genes to enhance the interpretability of gene variants.

Humans

[Variants of chronic heart failure in ischemic heart disease patients and optimization of their treatment].

Computer-assisted classification of hemodynamic data was performed in 172 patients with coronary heart disease aggravated by chronic heart failure. Six groups of patients have been identified, and an individual treatment algorithm has been proposed for each of those. The use of optimum individual treatment schedules has produced good or satisfactory clinical effect in 87.3%.

Adult

Analysis of adenomatous structures in histopathology.

A new idea of structure analysis in histopathology based upon first-order and third-order structures is presented. Networks formed by single cells and by tubulopapillary formations in adenomatous tissue were analyzed. The algorithm applied is based on the neighborhood conditions defined by O'Callaghan, using graph theory procedures. Twenty cases each of healthy colon mucosa, tubulovillous adenomas and highly to moderately differentiated adenocarcinomas of colon plus ten cases of mesotheliomas and ten cases of adenocarcinomas metastatic to the pleura were analyzed. Statistically significant differences were found in the cyclomatic number of neighboring elements. Classification of specimens of colon mucosa using discriminant analysis yielded correct results in 85% of the 20 cases. All ten cases of metastatic adenocarcinoma and nine of the ten cases of mesothelioma were also correctly classified by the same procedure. A trial of prospective diagnostic assistance in routine histology based upon these cases gave correct classification of three mesotheliomas and of two adenocarcinomas. The procedures are now being used successfully in the routine diagnosis of pleural epithelial/biphasic mesothelioma and of pleuritis carcinomatosa.

Adenocarcinoma