PubMed HealthSearch

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Automatic wave form classification of extracellular multineuron recordings.

A PC-based method for the reconstruction of individual spike trains from extracellular multineuron recordings is described. Starting with virtually no knowledge about the wave forms in a record, a fully automatic template-finding algorithm extracts templates using the entire data set. In a second step, individual spike trains are reconstructed.

Algorithms

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans

Comparison of the ID3 algorithm versus discriminant analysis for performing feature selection.

Having obtained disappointing results in a small medical data set despite the fact that our data seemed to be well suited for induction via ID3, we decided to compare the performance of ID3 to discriminant analysis. Performance was gauged by the percentage of correct classification in a second, independent data set. Examples were obtained from a cardiology project on the accuracy of auscultation. There were 107 examples in the first data set and 67 cases in the second. We found that ID3 and discriminant analysis performed equally poorly, with ID3 classifying only 60% of the second set correctly and discriminant analysis classifying 66% of the second set correctly. Also, the ID3 probability statistic for estimating the accuracy of ID3 for classifying further cases was markedly optimistic compared to our actual second data set results. Moreover, with an increase in sample size, ID3 seemed to break down, producing a large, complex decision tree of dubious generality, whereas discriminant analysis, with a larger sample size, used more independent variables but maintained its first set accuracy. These data suggest that there is a need for more sophisticated algorithms than ID3, even at the risk of giving up some computational efficiency.

Adult

Unsupervised clustering and centroid estimation using dynamic competitive learning.

In this paper, an unsupervised learning algorithm is developed. Two versions of an artificial neural network, termed a differentiator, are described. It is shown that our algorithm is a dynamic variation of the competitive learning found in most unsupervised learning systems. These systems are frequently used for solving certain pattern recognition tasks such as pattern classification and k-means clustering. Using computer simulation, it is shown that dynamic competitive learning outperforms simple competitive learning methods in solving cluster detection and centroid estimation problems. The simulation results demonstrate that high quality clusters are detected by our method in a short training time. Either a distortion function or the minimum spanning tree method of clustering is used to verify the clustering results. By taking full advantage of all the information presented in the course of training in the differentiator, we demonstrate a powerful adaptive system capable of learning continuously changing patterns.

Cluster Analysis

Identification of the elastic symmetry of bone and other materials.

A simplified classification scheme for the elastic symmetries of a solid is applied to the identification of the elastic symmetry of a material by three different methods--visual, stereological and numerical algorithm. Each method is illustrated with an application to bone tissues, but the methods apply to all materials.

Algorithms

HCSeeker: A classification tool for human genetic variant hot and cold spots designed for PM1 and benign criteria in the ACMG-AMP guideline.

PURPOSE: The PM1 criterion, which states that a variant is located in a mutational hot spot and/or critical and well-established functional domain without benign variation (such as the active site of an enzyme), is considered moderate evidence for assessing its pathogenicity. Although guidelines from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology are widely adopted, the PM1 criterion remains limited from lacking a reliable database of variant hot spots. Compared with hot spots, cold spots are neglected by the guidelines. To improve variant classification, we suggest including cold spots for supporting benign classifications. Consequently, we have developed the HCSeeker to provide data support for PM1 and the "Benign" criteria. METHODS: HCSeeker uses the Kernel Density Estimation and the Expectation-Maximization algorithm to identify hot- and cold-spot regions. RESULTS: Through HCSeeker, we identified 988 hot spots and 682 cold spots across 889 genes and provided a public database (http://www.genemed.tech/hcseeker/) for researchers and clinicians to query variant locations, facilitating the application of American College of Medical Genetics and Genomics and the Association for Molecular Pathology PM1 or "Benign" criteria. CONCLUSION: We developed the HCSeeker tool, which can effectively identify variant hot and cold spots within genes to enhance the interpretability of gene variants.

Humans

[Variants of chronic heart failure in ischemic heart disease patients and optimization of their treatment].

Computer-assisted classification of hemodynamic data was performed in 172 patients with coronary heart disease aggravated by chronic heart failure. Six groups of patients have been identified, and an individual treatment algorithm has been proposed for each of those. The use of optimum individual treatment schedules has produced good or satisfactory clinical effect in 87.3%.

Adult

Analysis of adenomatous structures in histopathology.

A new idea of structure analysis in histopathology based upon first-order and third-order structures is presented. Networks formed by single cells and by tubulopapillary formations in adenomatous tissue were analyzed. The algorithm applied is based on the neighborhood conditions defined by O'Callaghan, using graph theory procedures. Twenty cases each of healthy colon mucosa, tubulovillous adenomas and highly to moderately differentiated adenocarcinomas of colon plus ten cases of mesotheliomas and ten cases of adenocarcinomas metastatic to the pleura were analyzed. Statistically significant differences were found in the cyclomatic number of neighboring elements. Classification of specimens of colon mucosa using discriminant analysis yielded correct results in 85% of the 20 cases. All ten cases of metastatic adenocarcinoma and nine of the ten cases of mesothelioma were also correctly classified by the same procedure. A trial of prospective diagnostic assistance in routine histology based upon these cases gave correct classification of three mesotheliomas and of two adenocarcinomas. The procedures are now being used successfully in the routine diagnosis of pleural epithelial/biphasic mesothelioma and of pleuritis carcinomatosa.

Adenocarcinoma

Setting up an HIV screening program.

Human immunodeficiency virus-1 (HIV-1) screening programs currently are based primarily on the detection of specific HIV-1 antibodies by the commercially available enzyme immunoassay (EIA) combined with highly specific confirmation procedures. Factors to be considered in establishing a screening program include test performance characteristics, economy, confidentiality and notification procedures, legal and regulatory issues, proficiency and quality control measures, and laboratory safety. Commercial EIA screening in conjunction with a licensed Western blot assay permits the classification of all but a few serum samples into HIV-1-positive and HIV-1-negative categories. The occasional indeterminate results often can be resolved by following a defined retesting/resampling algorithm or by using research-level test procedures that may become available for diagnostic use in the future. Although screening of patient populations with an increased risk of HIV-1 exposure will improve the predictive accuracy of an initial screening assay, confirmation testing should nonetheless be performed for all EIA reactive sera regardless of the source. Local HIV-1 screening programs that meet minimum-volume requirements can result in considerable savings and flexibility for a moderate-size institution. However, before this type of program is undertaken, numerous technical and ethical considerations need to be addressed.

Algorithms

CSGL: chemical synthesis graph learning for molecule representation.

MOTIVATION: Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. RESULTS: Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-2023/CSGL.

Machine Learning

Prediction of spike-wave bursts in absence epilepsy by EEG power-spectrum signals.

The EEGs of subjects with absence seizures were examined to determine if changes occurred prior to spike-wave bursts that could be used to predict bursts. A number of 20-s epochs of EEG prior to spike-wave bursts (preburst epochs) and during periods remote from bursts (control epochs) were examined in 5 subjects. Power-spectrum analysis was carried out on each epoch and frequency bands from 0 to 50 c/s were combined into 2-c/s bandwidths. Logarithmically transformed power values in each frequency band were entered into a discriminant analysis algorithm for each subject separately. Results were expressed in terms of a test for significant differences between preburst and control epochs (F statistic) and a "success ratio" of discriminant analysis classification, defined as the proportion of correct classifications in both groups, as obtained using a cross-validation procedure. A significant preburst EEG pattern was found in 4 of the 5 subjects, and success ratios ranged from 0.64. to 0.83. Each subject's preburst EEG seemed to be characterized by a unique pattern of changes, and thus no common prodromal signal was found. The EEG changes did not appear to be caused by overt behaviors, such as eye closure or drowsiness. The findings suggest that the preburst EEG pattern represents a functional alteration in brain activity which could arise from the burst-producing mechanism directly.

Adolescent

Deciphering the Role of LNX2 as a Potential Contributor to Neurodevelopmental Disorders.

BACKGROUND/OBJECTIVES: Attention-deficit/hyperactivity disorder (ADHD) is a common neurodevelopmental condition characterized by a complex and multifactorial genetic architecture. In this study, we report a male patient, born to non-consanguineous healthy parents, presenting with ADHD and oppositional defiant disorder (ODD). METHODS: Trio-based whole-exome sequencing (WES) was performed in the proband and both parents. Variant classification was performed according to American College of Medical Genetics and Genomics (ACMG) guidelines, and the potential pathogenicity of the identified variant was further assessed through multiple in silico prediction algorithms and protein structural analyses. RESULTS: WES identified a homozygous variant in the LNX2 gene (NM_153371.4: c.1165G>A, p.Ala389Thr), classified as a variant of uncertain significance (VUS) and supported by multiple in silico predictions. LNX2 is expressed during brain development and encodes an E3 ubiquitin ligase involved in neuronal differentiation and synaptic function. The identified variant is located within the PDZ2 domain, a functionally relevant region involved in protein-protein interactions. Although the variant is reported in population databases (gnomAD ID: rs148429804), it has not been associated with any clinical phenotype, and its presence in the homozygous state has been reported only once, remaining extremely rare and lacking clinical annotation. Structural modelling predicted localized rearrangement of the hydrogen-bonding network within the PDZ2 domain without major conformational changes. Integrative transcriptomic, and single-cell analyses further supported the biological relevance of LNX2 in neurodevelopment, highlighting its preferential association with neuronal projection-cell networks, synaptic vesicle trafficking pathways, and neuron-specific regulatory programs. CONCLUSION: Although the identified LNX2 variant cannot be considered causative for the patient's phenotype and a definitive disease-gene relationship cannot be established based on a single individual, the complementary genetic, structural, and transcriptomic findings support the biological plausibility of LNX2 as a candidate gene for neurodevelopmental disorders. Additional independent patients and functional studies will be required to clarify its contribution to human disease.

Child

Lesions of the triangular fibrocartilaginous complex.

The authors present a retrospective study of 14 triangular fibrocartilaginous lesions in 13 patients. Based on their radiographic and clinical observations a new classification is presented taking into account associated lesions in the wrist.

Adolescent

NOVACODE serial ECG classification system for clinical trials and epidemiologic studies.

Traditional serial electrocardiogram (ECG) change classification schemes used in clinical trials such as the Minnesota Code rely on independent classification of the baseline and each follow-up or acute event ECG, whereby graded changes in the hierarchic severity level of the code signify new events such as myocardial infarction (MI). This approach suffers from classification errors caused by repeated instability at decision boundaries at each step when the baseline and each acute event ECG is classified, and various "verification rules" must be used at the end of the coding process to prevent trivial serial changes from causing large transitions in coded events. The NOVACODE algorithms for visual and computer coding of serial ECGs were designed to alleviate some of these instability problems by quantifying changes in critical waveform patterns on a continuous scale. This is achieved by determining, for each ECG coded, a Q-QS Score, ST Depression Score, ST Elevation Score and T-Wave Score, each ranging from 0 to 50. In the next step, a score is derived for ST-T evolution and this ST-T Evolution Score together with changes in Q-QS Score define criteria for a hierarchic mutually exclusive serial ECG change classification scheme that includes coding categories for Q-wave and non-Q wave MIs, equivocal Q wave evolution, evolving ischemic ST-T abnormalities, and various combinations of nonevolving Q-QS wave, and ST-T abnormalities.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

Ventricular beat classifier using fractal number clustering.

A two-stage ventricular beat 'associative' classification procedure is described. The first stage separates typical beats from extrasystoles on the basis of area and polarity rules. At the second stage, the extrasystoles are classified in self-organised cluster formations of adjacent shape parameter values. This approach avoids the use of threshold values for discrimination between ectopic beats of different shapes, which could be critical in borderline cases. A pattern shape feature conventionally called a 'fractal number', in combination with a polarity attribute, was found to be a good criterion for waveform evaluation. An additional advantage of this pattern classification method is its good computational efficiency, which affords the opportunity to implement it in real-time systems.

Algorithms

Topological maps of protein sequences.

A new method based on neural networks to cluster proteins into families is described. The network is trained with the Kohonen unsupervised learning algorithm, using matrix pattern representations of the protein sequences as inputs. The components (x, y) of these 20 x 20 matrix patterns are the normalized frequencies of all pairs xy of amino acids in each sequence. We investigate the influence of different learning parameters in the final topological maps obtained with a learning set of ten proteins belonging to three established families. In all cases, except in those where the synaptic vectors remains nearly unchanged during learning, the ten proteins are correctly classified into the expected families. The classification by the trained network of mutated or incomplete sequences of the learned proteins is also analysed. The neural network gives a correct classification for a sequence mutated in 21.5% +/- 7% of its amino acids and for fragments representing 7.5% +/- 3% of the original sequence. Similar results were obtained with a learning set of 32 proteins belonging to 15 families. These results show that a neural network can be trained following the Kohonen algorithm to obtain topological maps of protein sequences, where related proteins are finally associated to the same winner neuron or to neighboring ones, and that the trained network can be applied to rapidly classify new sequences. This approach opens new possibilities to find rapid and efficient algorithms to organize and search for homologies in the whole protein database.

Amino Acid Sequence

Numerical classification of Streptomyces and related genera.

Four hundred and seventy-five strains, which included 394 type cultures of Streptomyces and representatives of 14 other actinomycete genera, were studied. Overall similarities of these strains for 139 unit characters were determined by the SSM and SJ coefficients and clustering by the UPGMA algorithm. Test error and overlap between the phena defined were within acceptable limits. Cluster-groups were defined by the SSM coefficient at the 70.1% similarity (S) level and by the SJ coefficient at the 50% S-level. Clusters were distinguished at the 77.5% SSM and 63% SJ S-levels. Groupings obtained with the two coefficients were generally similar, but there were some changes in the definition and membership of cluster-groups and clusters. The phenetic data obtained, together with those from previous diverse studies, indicated that the genera Actinopycnidium, Actinosporangium, Chainia, Elytrosporangium, Kitasatoa and Microellobosporia should be reduced to synonyms of Streptomyces, while Intrasporangium, Nocardioides and Streptoverticillium remained as distinct genera in the family Streptomycetaceae. Nocardiopsis dassonvillei also showed strong phenetic affinity to Streptomyces, despite its chemotaxonomic differences. Actinomadura sensu stricto was phenetically distinguishable from Streptomyces and 'Nocardia' mediterranea was recognized as a taxon distinct from both these genera and from Nocardia sensu stricto. Most of the Streptomyces type cultures fell into one large cluster-group. At the 77.5% SSM S-level, they were recovered in 19 major and 40 minor clusters, with 18 strains recovered as single member clusters. The status of the latter as species was therefore confirmed. Most of the minor clusters, consisting of two to five strains, can also be regarded as species. The major clusters varied in size (from 6 to 71 strains) and in there homogeneity. Therefore, it is suggested that they be regarded as species-groups until further information is available. The results provide a basis for the reduction of the large number of Streptomyces species which have been described. They also demonstrate that the previous use of a limited number of subjectively chosen characters to define species-groups or species has resulted in artificial classifications.

Culture Media

Rough sets approach to analysis of data from peritoneal lavage in acute pancreatitis.

Two kinds of information systems composed of data from peritoneal lavage in acute pancreatitis are analysed with the concept of rough sets: system A, classifying patients described by pre-lavage attributes, and system B, classifying patients described by attributes of the course of the multistage lavage. The analysis tends to define subsets which are significant for high quality of classification. These attributes give the best description of the patient's state and are in the closest relationship with the time of lavage. The character of this relationship is shown by decision algorithms derived from decision tables representing the information system.

Acute Disease