PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Real-time computing without stable states: a new framework for neural computation based on perturbations.

A key challenge for neural modeling is to explain how a continuous stream of multimodal input from a rapidly changing environment can be processed by stereotypical recurrent circuits of integrate-and-fire neurons in real time. We propose a new computational model for real-time computing on time-varying input that provides an alternative to paradigms based on Turing machines or attractor neural networks. It does not require a task-dependent construction of neural circuits. Instead, it is based on principles of high-dimensional dynamical systems in combination with statistical learning theory and can be implemented on generic evolved or found recurrent circuitry. It is shown that the inherent transient dynamics of the high-dimensional dynamical system formed by a sufficiently large and heterogeneous neural circuit may serve as universal analog fading memory. Readout neurons can learn to extract in real time from the current state of such recurrent neural circuit information about current and past inputs that may be needed for diverse tasks. Stable internal states are not required for giving a stable output, since transient internal states can be transformed by readout neurons into stable target outputs due to the high dimensionality of the dynamical system. Our approach is based on a rigorous computational model, the liquid state machine, that, unlike Turing machines, does not require sequential transitions between well-defined discrete internal states. It is supported, as the Turing machine is, by rigorous mathematical results that predict universal computational power under idealized conditions, but for the biologically more realistic scenario of real-time processing of time-varying inputs. Our approach provides new perspectives for the interpretation of neural coding, the design of experiments and data analysis in neurophysiology, and the solution of problems in robotics and neurotechnology.

Action Potentials↗

A generalized hidden Markov model for the recognition of human genes in DNA.

We present a statistical model of genes in DNA. A Generalized Hidden Markov Model (GHMM) provides the framework for describing the grammar of a legal parse of a DNA sequence (Stormo & Haussler 1994). Probabilities are assigned to transitions between states in the GHMM and to the generation of each nucleotide base given a particular state. Machine learning techniques are applied to optimize these probabilities using a standardized training set. Given a new candidate sequence, the best parse is deduced from the model using a dynamic programming algorithm to identify the path through the model with maximum probability. The GHMM is flexible and modular, so new sensors and additional states can be inserted easily. In addition, it provides simple solutions for integrating cardinality constraints, reading frame constraints, "indels", and homology searching. The description and results of an implementation of such a gene-finding model, called Genie, is presented. The exon sensor is a codon frequency model conditioned on windowed nucleotide frequency and the preceding codon. Two neural networks are used, as in (Brunak, Engelbrecht, & Knudsen 1991), for splice site prediction. We show that this simple model performs quite well. For a cross-validated standard test set of 304 genes [ftp:@www-hgc.lbl.gov/pub/genesets] in human DNA, our gene-finding system identified up to 85% of protein-coding bases correctly with a specificity of 80%. 58% of exons were exactly identified with a specificity of 51%. Genie is shown to perform favorably compared with several other gene-finding systems.

Chromosomes, Human↗

Decision tree-based formation of consensus protein secondary structure prediction.

MOTIVATION: Prediction of protein secondary structure provides information that is useful for other prediction methods like fold recognition and ab initio 3D prediction. A consensus prediction constructed from the output of several methods should yield more reliable results than each of the individual methods. METHOD: We present an approach that reveals subtle but systematic differences in the output of different secondary structure prediction methods allowing the derivation of coherent consensus predictions. The method uses a machine learning technique that builds decision trees from existing data. RESULTS: The first results of our analysis show that consensus prediction of protein secondary structure may be improved both quantitatively and qualitatively.

Algorithms↗

Proteomics-enabled learning machine algorithms enhance the prediction of cardiovascular diseases in patients with type 2 diabetes mellitus.

BACKGROUND AND AIMS: Estimating the risk of cardiovascular disease (CVD) complications in type 2 diabetes mellitus (T2DM) patients is critical in the medical decision-making process. This study aimed to use a machine learning technique combined with proteomics to develop personalized models for predicting CVD in patients with T2DM. METHODS AND RESULTS: In total, 874 patients with T2DM and 2,920 Olink proteins obtained from the UK Biobank were used in this study. Proteins were screened using Cox regression and LASSO regression. A basic model containing clinical features and a full model combining proteome and clinical features were constructed using the random survival forest algorithm. The area under the receiver operating characteristic (ROC) curve (AUC) was used to evaluate the predictive performance of the models and compare them with other CVD predictive models. Compared with the basic model, the full model performed better in predicting CVD, with time-dependent AUCs of 0.81 (3 years), 0.74 (5 years) and 0.74 (10 years) (0.77, 0.69 and 0.67). We calculated the risk scores of the Framingham, ASCVD and Score2-Diabetes models. The results revealed that the prediction performance of the full model was also better than that of the abovementioned models. In terms of differentiation accuracy, the results of the net reclassification improvement index and integrated discrimination improvement index showed that the full model can identify high-risk individuals more accurately (accuracy rate: 79% vs. 69%). CONCLUSIONS: Proteomics can be used to predict cardiovascular complications in diabetic patients. It is also necessary to consider the applicability of the model due to the limitations of the sample size and the constraints of proteomics in clinical applications.

Humans↗

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.

Large Language Models↗

Integrating single-cell transcriptomics to construct an oncogene-driven prognostic model and elucidate metabolic-immune crosstalk in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is a leading cause of cancer-related deaths, its progression and treatment heterogeneity are mainly influenced by driver gene and tumor micro-environment (TME) interactions. Nevertheless, the mechanisms of this process at the single-cell level remain unclear. This study integrated TCGA and multi-center single-cell transcriptome data to identify a 575 genes HCC-specific core set, developing a single-cell "oncogene scoring" system to quantify individual carcinogenic activity. This score is significantly elevated in malignant and proliferative T cells and is closely associated with metabolic reprogramming, aberrant cell‒cell communication, and immunosuppressive phenotypes. Based on these characteristics, we constructed a machine learning-based Random Survival Forest (RSF) prognostic model validated in multiple independent cohorts, which classifies patients into distinct risk subtypes. The high-risk group exhibits genomic instability, increased tumor stemness, and immune evasion, while the low-risk group was more sensitive to drugs such as sorafenib. This study highlights the potential pathways by which high oncogenic activity is associated with HCC progression, suggesting a profound link with single-cell metabolic‒immune crosstalk. The constructed RSF model offers a promising computational framework for risk stratification and provides hypothesis-generating insights that may inform future personalized treatment strategies for HCC patients.

Hepatocellular carcinoma↗

Osteoarthritis phenotypes: advancing precision medicine through clinical, structural, and molecular stratification.

PURPOSE: Osteoarthritis (OA) is now understood as a heterogeneous syndrome driven by diverse biological, biomechanical, metabolic, genetic, and molecular mechanisms. This variability explains differences in disease progression and treatment response, challenging the traditional "one-size-fits-all" approach. This review highlights OA phenotyping as a key step toward precision medicine, focusing on clinical, structural, and molecular classifications that inform individualized care. METHODS: A narrative review was conducted using a non-systematic search of major databases and Osteoarthritis Research Society International sources (2010-2026). Evidence was thematically synthesized across clinical, imaging, and molecular domains to characterize OA phenotypes and their potential relevance to precision medicine. RESULTS: Multiple OA phenotypes were identified: inflammatory, metabolic, biomechanical, cartilage-subchondral, pain-sensitization, and aging/senescence. These exhibit distinct clinical features, risk factors, and therapeutic responses. Imaging-based phenotypes (e.g., inflammatory, meniscus-cartilage, subchondral bone, atrophic, hypertrophic) and molecular endotypes (low turnover, structural damage, systemic inflammation) further refine stratification. Pain-structure discordance is notable in sensitization phenotypes and may predict poorer surgical outcomes. Joint-specific variations and emerging genomic and epigenetic insights underscore disease complexity. Advances in imaging, biomarkers, and machine learning may enable earlier detection and patient clustering, though clinical application remains limited. CONCLUSION: Phenotype- and endotype-based classification represents a critical advancement toward precision OA management. Tailored interventions based on stratification hold promise for improving outcomes; however, clinical translation remains limited by overlapping phenotypes, lack of validated biomarkers, and inconsistent results from phenotype-driven trials. Wider clinical adoption requires standardized definitions, validation across joints, and integration of multimodal diagnostic tools into routine practice.

Humans↗

Automatic synthesis of synergies for control of reaching--hierarchical clustering.

In this paper we describe a novel method for determining synergies between joint motions in reaching movements by hierarchical clustering. A set of recorded elbow and shoulder trajectories is used in a learning algorithm to determine the relationships between angular velocities at elbow and shoulder joints. The learning algorithm is based on optimal criteria for obtaining the hierarchy of descriptions of movement trajectories. We show that this method finds complex synergism between optimal joint trajectories for a given set of data and angular velocities at the shoulder and elbow joints. Three other machine learning techniques (ML) are used for comparison with our method of hierarchical clustering of trajectories. These MLs are: (1) radial basis functions (RBF), (2) inductive learning (IL), and (3) adaptive-network-based fuzzy inference system (ANFIS). Better error characteristics were obtained using the method of hierarchical clustering in comparison with the other techniques. The advantage of the method of hierarchical clustering with respect to the other MLs is in integrating the spatial and temporal elements of reaching movements. Determination and analysis of spatio-temporal events of movement trajectories is a useful tool in designing control systems for functional electrical stimulation (FES) assisted manipulation.

Algorithms↗

CCNA2 orchestrates the PI3K/AKT signaling axis to propel prostate cancer metastasis.

BACKGROUND: Prostate cancer (PCa) remains one of the most common malignancies in men, posing a persistent global burden in terms of both public health and socioeconomic costs. Although early detection is essential for improving patient outcomes, existing clinical tools, including prostate-specific antigen (PSA) screening, digital rectal examination, and transrectal ultrasound-guided biopsy, are hampered by suboptimal specificity and positive predictive value, resulting in frequent overdiagnosis and overtreatment of indolent lesions while missing a subset of aggressive tumors at an early stage. In this context, the rapid advancement of high-throughput omics technologies, coupled with sophisticated machine learning (ML) algorithms, provides a powerful computational framework to dissect high-dimensional genomic data, uncover latent gene expression signatures, and identify candidate biomarkers with superior discriminative performance over conventional clinicopathological parameters. Therefore, in this study, we sought to screen for crucial ML-based biomarkers associated with PCa, with a particular focus on systematically assessing the diagnostic and prognostic value of CCNA2. Leveraging large-scale transcriptomic cohorts from public repositories, we employed an ensemble of ML approaches to prioritize candidate genes and subsequently evaluated the diagnostic performance of CCNA2 through receiver operating characteristic curve analysis, as well as its prognostic utility via Kaplan-Meier survival estimation and multivariate Cox proportional hazards modeling. Our findings are anticipated to elucidate the molecular landscape of PCa and offer a promising biomarker candidate for early detection and risk stratification. METHODS: This study integrated single-cell RNA sequencing, bulk transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories, immunofluorescence, and multiple ML algorithms with in vitro functional assays to evaluate CCNA2 expression, clinical relevance, and biological behavior in PCa. RESULTS: CCNA2 was linked to metastasis and poor prognosis. High CCNA2 expression significantly correlated with adverse survival outcomes, and knockdown of CCNA2 suppressed proliferation, migration, and invasion in PCa cell lines. Mechanistically, CCNA2 modulated the PI3K/AKT signaling pathway. An ML-based diagnostic model incorporating CCNA2 demonstrated high predictive accuracy across multiple validation cohorts. CONCLUSIONS: CCNA2 serves as a promising prognostic biomarker and therapeutic target in prostate adenocarcinoma, driving tumor progression potentially via the PI3K/AKT axis.

CCNA2↗

VR interaction techniques for medical imaging applications.

Methods of virtual reality (VR) offer new ways of human-computer interaction. Medicine is predestined to benefit from this new technology in many ways. Virtual environments can support physicians in their work, alleviate communication between specialists from different fields or be established in educational and training applications. For the field of visualization and analysis of three-dimensional anatomical images (e.g. CT or MRI scans), an application is introduced which expedites recognition of spatial coherencies and the exploration and manipulation of the 3D data. To avoid long periods of learning and accustoming and to facilitate work in such an environment, a powerful human-oriented interface is required allowing interactions similar to the real world and utilization of our natural experiences. This paper shows the use of eye tracking parameters for a level-of-detail algorithm and the integration of a glove-based hand gesture recognition into the virtual environment as an essential component of the human-machine interface. Furthermore, virtual bronchoscopy and virtual angioscopy are presented as examples for the use of the virtual environment.

Diagnosis, Computer-Assisted↗

Nanocarrier-Based Gene Delivery Systems: Mechanisms, Clinical Translation, and Future Perspectives.

Gene therapy holds revolutionary potential for managing genetic disorders, cancers and infectious illnesses. However, one of the biggest challenges is delivering DNA or RNA into targeted cells and in the safe and effective way. In this review, nano carrier-based approaches for gene delivery are critically examined, focusing on both viral and non-viral systems. The advancement of CRISPR-Cas genome editing, machine learning-assisted nanocarrier optimization, and biologically inspired delivery systems is being quickly pushed forward in this area. In this review, a comparative analysis of gene delivery systems is being provided, and the key challenges to clinical translation are being pointed out. In addition, expert opinions on future research directions are being offered, with a heavy focus on the development of multifunctional, precisely targeted, and easily scalable delivery systems that can be integrated with next-generation therapeutic technologies.

Humans↗

DNA methylation biomarkers for early detection of ovarian cancer.

Ovarian cancer (OC) remains difficult to detect at an early stage, and current screening approaches using CA125 and transvaginal ultrasonography have not demonstrated sufficient benefit for population screening. DNA methylation is a promising biomarker class because epigenetic alterations may arise early in tumourigenesis, can be detected in circulating cell-free DNA (cfDNA), and may provide tissue-of-origin information. This review critically evaluates recent evidence on DNA methylation biomarkers for early OC detection. PubMed/MEDLINE, Web of Science, and Scopus were searched for studies published between January 2020 and September 2025, supplemented by selected earlier studies of biological or methodological relevance. Evidence was synthesised across single-gene biomarkers, multi-locus panels, genome-wide signatures, assay platforms, and machine-learning classifiers, with emphasis on early-stage performance, histological representation, comparator populations, analytical methodology, and validation design. Single-gene markers such as BRCA1, RASSF1A, OPCML, HOXA9, and HIC1 show variable performance, while multi-gene and classifier-based approaches generally provide stronger discrimination. However, many studies remain limited by retrospective case-control designs, small FIGO stage I-II subsets, predominance of serous disease, and insufficient prospective validation. Integration with CA125 may improve sensitivity but can reduce specificity, which is critical in low-prevalence screening. Clinical translation will therefore require minimal and reproducible methylation signatures, standardised low-input cfDNA workflows, rigorous external validation, and prospective longitudinal evaluation in intended-use populations.

Humans↗

Computer-integrated revision total hip replacement surgery: concept and preliminary results.

This paper describes an ongoing project to develop a computer-integrated system to assist surgeons in revision total hip replacement (RTHR) surgery. In RTHR surgery, a failing orthopedic hip implant, typically cemented, is replaced with a new one by removing the old implant, removing the cement and fitting a new implant into an enlarged canal broached in the femur. RTHR surgery is a difficult procedure fraught with technical challenges and a high incidence of complications. The goals of the computer-based system are the significant reduction of cement removal labor and time, the elimination of cortical wall penetration and femur fracture, the improved positioning and fit of the new implant resulting from precise, high-quality canal milling and the reduction of bone sacrificed to fit the new implant. Our starting points are the ROBODOC system for primary hip replacement surgery and the manual RTHR surgical protocol. We first discuss the main difficulties of computer-integrated RTHR surgery and identify key issues and possible solutions. We then describe possible system architectures and protocols for preoperative planning and intraoperative execution. We present a summary of methods and preliminary results in CT image metal artifact removal, interactive cement cut-volume definition and cement machining, anatomy-based registration using fluoroscopic X-ray images and clinical trials using an extended RTHR version of ROBODOC. We conclude with a summary of lessons learned and a discussion of current and future work.

Algorithms↗

Dynamic probability estimator for machine learning.

An efficient algorithm for dynamic estimation of probabilities without division on unlimited number of input data is presented. The method estimates probabilities of the sampled data from the raw sample count, while keeping the total count value constant. Accuracy of the estimate depends on the counter size, rather than on the total number of data points. Estimator follows variations of the incoming data probability within a fixed window size, without explicit implementation of the windowing technique. Total design area is very small and all probabilities are estimated concurrently. Dynamic probability estimator was implemented using a programmable gate array from Xilinx. The performance of this implementation is evaluated in terms of the area efficiency and execution time. This method is suitable for the highly integrated design of artificial neural networks where a large number of dynamic probability estimators can work concurrently.

Artificial Intelligence↗

Integrative chemical genetics platform identifies condensate modulators linked to neurological disorders.

Dysregulation of biomolecular condensates is implicated across multiple neurological disorders. However, approaches to systematically identify their modulators remain limited. Here, we expand the utility of MLF2 as a versatile condensate biomarker and develop CondenScreen, an integrated high-content screening and bioinformatics pipeline enabling identification of condensate modulators across chemical and genetic space. Screening 1760 bioactive compounds in a cellular DYT1 dystonia model, we validate the platform for condensate-targeted drug discovery, identifying drugs that prevent the accumulation of the MLF2 reporter into nuclear envelope condensates. In parallel, a genome-wide CRISPR/Cas9 screen correlates nuclear condensate abundance with genes implicated in microcephaly and over eight additional neurodevelopmental disorders. Machine learning and confocal imaging resolve distinct condensate phenotypes, with RNF26 deletion provoking nuclear envelope condensates that phenocopy hallmarks of torsin deficiency. Our study provides a scalable platform for identifying modulators of condensates and establishes a correlative connection between nuclear condensate accumulation and genes implicated in neurodevelopmental disorders.

Humans↗

Uncovering the genetic architecture of ME/CFS: a precision approach reveals impact of rare monogenic variation.

BACKGROUND: Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a disabling and heterogeneous disorder lacking validated biomarkers or targeted therapies. Clinical variability and elusive pathophysiology hinder progress toward effective diagnostics and treatment. Core symptoms include persistent fatigue, post-exertional malaise, unrefreshing sleep, cognitive dysfunction, and pain. We tested whether an individualized, “n-of-1” genomic and transcriptomic framework combined with comprehensive, participant-informed phenotyping could reveal molecular signatures unique to each patient. METHODS: Clinical-grade whole-genome sequencing was conducted in 31 affected individuals from 25 families, with RNA-seq performed on a subset (16 affected, 7 unaffected) using blood samples. Machine-learning assisted variant triage, transcript-aware damage prediction, and expert review identified pathogenic or likely pathogenic variants in 8 of 25 probands (32%) and 12 of 31 affected individuals (39%). RESULTS: Findings revealed marked genetic heterogeneity, including large-effect rare and more common variants. Implicated pathways included ATP generation, oxidative phosphorylation, fatty acid oxidation; regulation of glycolysis, amino acid and lipid turnover; ion and solute homeostasis; synaptic signaling, excitability, oxygen transport, and muscle integrity, resilience, and post-exertional recovery; previously implicated processes. Plausible modifiers influencing disease onset, severity, and relapsing–remitting patterns and possibly explaining intrafamilial variability and inconsistent findings across studies, were also identified. Despite gene-level diversity, downstream effects converged on impaired energy production, reduced stress resilience, and vulnerability to post-exertional metabolic failure; disruptions consistent with core ME/CFS symptoms of exertional intolerance, cognitive fog, and fatigue. CONCLUSIONS: Our findings support the hypothesis that at least a subset of ME/CFS cases represent distinct molecular disorders that converge on shared physiological pathways. Validation in larger, more diverse cohorts will be essential to test this hypothesis and establish generalizability, but increase size alone is unlikely to resolve causation in a disorder defined by rarity, heterogeneity, and molecular complexity. We suggest that progress will require experimental designs that integrate individual-level genomic data with deep, participant-informed deep phenotyping, capturing the combined effects of rare and common variants and environmental modifiers on disease expression and progression. We believe that an individualized precision medicine framework will uncover molecular drivers and modifiers of ME/CFS previously obscured by heterogeneity, enabling biologically informed stratification, improved trial design, biomarker discovery, and targeted interventions in this historically neglected condition.

Humans↗

Molecular scene analysis: the integration of direct-methods and artificial-intelligence strategies for solving protein crystal structure.

A knowledge-based approach to crystal structure determination is presented. The approach integrates direct-methods and artificial-intelligence strategies to rephrase the structure determination process as an exercise in scene analysis. A general joint probability distribution framework, which allows the incorporation of isomorphous replacement, anomalous scattering and a priori structural information, forms the basis of the direct-methods strategies. The accumulated knowledge on crystal and molecular structures is exploited through the use of artificial-intelligence strategies, which include techniques of knowledge representation, search and machine learning.

Journal Article↗

An incremental approach to genetic-algorithms-based classification.

Incremental learning has been widely addressed in the machine learning literature to cope with learning tasks where the learning environment is ever changing or training samples become available over time. However, most research work explores incremental learning with statistical algorithms or neural networks, rather than evolutionary algorithms. The work in this paper employs genetic algorithms (GAs) as basic learning algorithms for incremental learning within one or more classifier agents in a multiagent environment. Four new approaches with different initialization schemes are proposed. They keep the old solutions and use an "integration" operation to integrate them with new elements to accommodate new attributes, while biased mutation and crossover operations are adopted to further evolve a reinforced solution. The simulation results on benchmark classification data sets show that the proposed approaches can deal with the arrival of new input attributes and integrate them with the original input space. It is also shown that the proposed approaches can be successfully used for incremental learning and improve classification rates as compared to the retraining GA. Possible applications for continuous incremental training and feature selection are also discussed.

Algorithms↗