PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Acetyl-coenzyme A synthetase (AMP forming).

Acetyl-coenzyme A synthetase (AMP forming; Acs) is an enzyme whose activity is central to the metabolism of prokaryotic and eukaryotic cells. The physiological role of this enzyme is to activate acetate to acetyl-coenzyme A (Ac-CoA). The importance of Acs has been recognized for decades, since it provides the cell the two-carbon metabolite used in many anabolic and energy generation processes. In the last decade researchers have learned how carefully the cell monitors the synthesis and activity of this enzyme. In eukaryotes and prokaryotes, complex regulatory systems control acs gene expression as a function carbon flux, with a second layer of regulation exerted posttranslationally by the NAD+/sirtuin-dependent protein acetylation/deacetylation system. Recent structural work provides snapshots of the dramatic conformational changes Acs undergoes during catalysis. Future work on the regulation of acs gene expression will expand our understanding of metabolic integration, while structure/function studies will reveal more details of the function of this splendid molecular machine.

Acetate-CoA Ligase↗

Machine learning approaches for the prediction of signal peptides and other protein sorting signals.

Prediction of protein sorting signals from the sequence of amino acids has great importance in the field of proteomics today. Recently, the growth of protein databases, combined with machine learning approaches, such as neural networks and hidden Markov models, have made it possible to achieve a level of reliability where practical use in, for example automatic database annotation is feasible. In this review, we concentrate on the present status and future perspectives of SignalP, our neural network-based method for prediction of the most well-known sorting signal: the secretory signal peptide. We discuss the problems associated with the use of SignalP on genomic sequences, showing that signal peptide prediction will improve further if integrated with predictions of start codons and transmembrane helices. As a step towards this goal, a hidden Markov model version of SignalP has been developed, making it possible to discriminate between cleaved signal peptides and uncleaved signal anchors. Furthermore, we show how SignalP can be used to characterize putative signal peptides from an archaeon, Methanococcus jannaschii. Finally, we briefly review a few methods for predicting other protein sorting signals and discuss the future of protein sorting prediction in general.

Algorithms↗

LS-NMF: a modified non-negative matrix factorization algorithm utilizing uncertainty estimates.

BACKGROUND: Non-negative matrix factorisation (NMF), a machine learning algorithm, has been applied to the analysis of microarray data. A key feature of NMF is the ability to identify patterns that together explain the data as a linear combination of expression signatures. Microarray data generally includes individual estimates of uncertainty for each gene in each condition, however NMF does not exploit this information. Previous work has shown that such uncertainties can be extremely valuable for pattern recognition. RESULTS: We have created a new algorithm, least squares non-negative matrix factorization, LS-NMF, which integrates uncertainty measurements of gene expression data into NMF updating rules. While the LS-NMF algorithm maintains the advantages of original NMF algorithm, such as easy implementation and a guaranteed locally optimal solution, the performance in terms of linking functionally related genes has been improved. LS-NMF exceeds NMF significantly in terms of identifying functionally related genes as determined from annotations in the MIPS database. CONCLUSION: Uncertainty measurements on gene expression data provide valuable information for data analysis, and use of this information in the LS-NMF algorithm significantly improves the power of the NMF technique.

Algorithms↗

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans↗

miRNAMap: genomic maps of microRNA genes and their target genes in mammalian genomes.

Recent work has demonstrated that microRNAs (miRNAs) are involved in critical biological processes by suppressing the translation of coding genes. This work develops an integrated database, miRNAMap, to store the known miRNA genes, the putative miRNA genes, the known miRNA targets and the putative miRNA targets. The known miRNA genes in four mammalian genomes such as human, mouse, rat and dog are obtained from miRBase, and experimentally validated miRNA targets are identified in a survey of the literature. Putative miRNA precursors were identified by RNAz, which is a non-coding RNA prediction tool based on comparative sequence analysis. The mature miRNA of the putative miRNA genes is accurately determined using a machine learning approach, mmiRNA. Then, miRanda was applied to predict the miRNA targets within the conserved regions in 3'-UTR of the genes in the four mammalian genomes. The miRNAMap also provides the expression profiles of the known miRNAs, cross-species comparisons, gene annotations and cross-links to other biological databases. Both textual and graphical web interface are provided to facilitate the retrieval of data from the miRNAMap. The database is freely available at http://mirnamap.mbc.nctu.edu.tw/.

Animals↗

The application of artificial intelligence in healthcare practice: A mapping review of systematic reviews.

Artificial intelligence (AI) is rapidly transforming healthcare practice, with growing evidence supporting its use in diagnosis, prognosis, treatment planning, and operational decision-making. The proliferation of systematic reviews in recent years underscores the need for an updated synthesis of the literature to inform research, policy, and practice. We searched PubMed, Web of Science, Scopus, IEEE Xplore, and CINAHL for systematic reviews and meta-analyses published between 2019 and February 2026. Eligible reviews focused on AI applications in healthcare practice, were peer-reviewed, and written in English. A total of 368 reviews met the inclusion criteria. Publication volume increased steadily, peaking in 2025. AI research was concentrated in high-density domains, such as radiology, oncology, and critical care. Across reviews, diagnostic imaging, electronic health record (EHR) data, and biomarkers/laboratory results accounted for 68% of training data sources, though newer data types, such as wearable device and sensor data, emerged from 2022 onward. Diagnosis, prognosis, and treatment comprised over 80% of AI applications, with novel uses emerging in recent years, such as AI-assisted clinical documentation (e.g., ambient documentation tools) and patient education. Ethical concerns were reported in 78.5% of reviews, with privacy, model accuracy, data and algorithmic bias, and explainability as recurrent themes. The proportion of reviews reporting ethical concerns increased from 2021 to 2025. AI applications in healthcare are expanding in scope, diversifying in data sources, and evolving toward novel clinical and operational uses. The human-centered AI or augmented intelligence paradigm, integrating computational precision with clinical expertise, holds significant promise but will require parallel advances in governance, regulatory frameworks, and ethical oversight to ensure safe adoption.

Artificial Intelligence↗

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar↗

SIGEL: a context-aware genomic representation learning framework for spatial genomics analysis.

Spatial transcriptomics (ST) integrates spatial information into genomics, yet methods for generating spatially-informed gene representations are limited and computationally intensive. We present SIGEL, a cost-effective framework that derives gene manifolds from ST data by exploiting spatial genomic context. The resulting SIGEL-generated gene representations (SGRs) are context-aware, biologically meaningful, and robust across samples, making them highly effective for key downstream tasks, including imputing missing genes, detecting spatial expression patterns, identifying disease-related genes and interactions, and improving spatial clustering. Extensive experiments across diverse ST datasets validate SIGEL's effectiveness and highlight its potential in advancing spatial genomics research.

Genomics↗

[A diagnostic expert system].

The introduction of new technologies in the field of electronics has influenced the development of technical equipment over the last few years. The progressive miniaturization of integrated circuits makes possible an expansion of the spectrum of functions offered by this equipment. This also applies to medical technology. These more complex units call for new methods of fault detection and diagnosis. In addition to analytical redundancy, tools developed by the artificial intelligence research community, such as expert systems, are becoming more and more important for fault diagnosis. On the basis of a realized diagnosis expert system the possibilities as well as the limits of such system are discussed. Also, possible future developments of artificial intelligence, like machine learning, are considered.

Artificial Intelligence↗

Comprehensive in silico genomics analysis of global trends and host-specific emergence of aminoglycoside resistance in Staphylococcus aureus: a One-Health perspective.

BACKGROUND: Aminoglycosides remain clinically valuable against Staphylococcus aureus. Aminoglycoside resistance in S. aureus represents a critical One Health concern and is primarily driven by aminoglycoside-modifying enzymes (AMEs), which are frequently plasmid-encoded. Although regional studies have provided valuable insights, the global epidemiology of aminoglycoside resistance determinants remains poorly characterized because comprehensive data integrating human, animal, and environmental reservoirs are still lacking. This study addresses this gap by analyzing over 110,000 S. aureus genomes (2000-2025) to map the global resistome, quantify temporal and host-specific trends, and assess the association between genetic determinants and phenotypic resistance. METHODS: We performed a retrospective One Health meta-analysis of 110,309 S. aureus genomes collected between 2000 and 2025 from 128 countries. Genomes were quality-filtered and aminoglycoside resistance determinants were identified using NCBI AMRFinderPlus (v4.0.23). Multilocus sequence typing and host-source harmonization (Human, Animal, Environment, Unknown) enabled clonal and reservoir stratification. Temporal trends in gene prevalence and resistance burden were modeled with robust regression. Geographic and host-associated structuring of key genes was assessed via &#x3c7;2 and enrichment tests. Machine-learning models (elastic-net, random forests, XGBoost) were benchmarked for minimum inhibitory concentration (MIC) prediction via nested cross-validation, with performance evaluated by mean absolute error, RMSE, and SHAP-based feature importance. All analyses were conducted in R and Python using publicly available, de-identified genomic data. RESULTS: Aminoglycoside resistance-associated genes were dominated by modifying enzyme determinants, with ant(6)-Ia, ant(9)-Ia, aph(3')-IIIa, sat4, aadD1, and aac(6')-Ie/aph(2'')-Ia occurring in 14-22% of isolates worldwide. Temporal analysis revealed significant declines in several major determinants, most notably ant(9)-Ia (-2.22 percentage points per year, p&#x2009;<&#x2009;0.001), whereas apmA exhibited a non-significant decreasing trend in animal isolates. Host structuring was marked: human clinical isolates concentrated common determinants, while animal and environmental isolates harbored rare alleles (apmA, spw, str, spd). Geographic mapping confirmed near-universal distribution of common genes but focal restriction of rare ones. Publicly available phenotypic data indicated strong activity of amikacin, whereas gentamicin showed a distinct resistant subpopulation that closely corresponded with AME gene carriage. Genotype-phenotype analyses demonstrated strong concordance, with gene-rich complements predicting resistant MIC strata and absence of determinants predicting susceptibility. Analysis across different gene classes revealed frequent co-occurrence of aminoglycoside resistance genes with determinants from other classes, such as mecA, blaZ, and MLS_B, embedding them within multidrug-resistant (MDR) genomic contexts. CONCLUSION: Over 25&#xa0;years, the prevalence of aminoglycoside resistance-associated genes in S. aureus has declined for several common determinants, while rare veterinary-linked alleles are emerging in animal isolates. Strong genotype-phenotype concordance supports genomic prediction for gentamicin and amikacin, where MIC data are available, although phenotypic confirmation remains essential. The frequent co-occurrence of aminoglycoside resistance genes with other antimicrobial resistance determinants indicates their integration within co-occurrence patterns of MDR genes, defined here as clusters of co-occurring resistance genes often carried on shared mobile genetic elements. These patterns highlight the need for integrated One Health surveillance combining clinical, veterinary, and environmental monitoring with plasmid-context resolution to anticipate emerging threats.

Aminoglycosides↗

Predicting genetic regulatory response using classification.

MOTIVATION: Studying gene regulatory mechanisms in simple model organisms through analysis of high-throughput genomic data has emerged as a central problem in computational biology. Most approaches in the literature have focused either on finding a few strong regulatory patterns or on learning descriptive models from training data. However, these approaches are not yet adequate for making accurate predictions about which genes will be up- or down-regulated in new or held-out experiments. By introducing a predictive methodology for this problem, we can use powerful tools from machine learning and assess the statistical significance of our predictions. RESULTS: We present a novel classification-based method for learning to predict gene regulatory response. Our approach is motivated by the hypothesis that in simple organisms such as Saccharomyces cerevisiae, we can learn a decision rule for predicting whether a gene is up- or down-regulated in a particular experiment based on (1) the presence of binding site subsequences ('motifs') in the gene's regulatory region and (2) the expression levels of regulators such as transcription factors in the experiment ('parents'). Thus, our learning task integrates two qualitatively different data sources: genome-wide cDNA microarray data across multiple perturbation and mutant experiments along with motif profile data from regulatory sequences. We convert the regression task of predicting real-valued gene expression measurements to a classification task of predicting +1 and -1 labels, corresponding to up- and down-regulation beyond the levels of biological and measurement noise in microarray measurements. The learning algorithm employed is boosting with a margin-based generalization of decision trees, alternating decision trees. This large-margin classifier is sufficiently flexible to allow complex logical functions, yet sufficiently simple to give insight into the combinatorial mechanisms of gene regulation. We observe encouraging prediction accuracy on experiments based on the Gasch S.cerevisiae dataset, and we show that we can accurately predict up- and down-regulation on held-out experiments. We also show how to extract significant regulators, motifs and motif-regulator pairs from the learned models for various stress responses. Our method thus provides predictive hypotheses, suggests biological experiments, and provides interpretable insight into the structure of genetic regulatory networks. AVAILABILITY: The MLJava package is available upon request to the authors. Supplementary: Additional results are available from http://www.cs.columbia.edu/compbio/geneclass

Binding Sites↗

Implementing tutoring strategies into a patient simulator for clinical reasoning learning.

OBJECTIVE: This paper describes an approach for developing intelligent tutoring systems (ITS) for teaching clinical reasoning. MATERIALS AND METHODS: Our approach to ITS for clinical reasoning uses a novel hybrid knowledge representation for the pedagogic model, combining finite state machines to model different phases in the diagnostic process, production rules to model triggering conditions for feedback in different phases, temporal logic to express triggering conditions based upon past states of the student's problem solving trace, and finite state machines to model feedback dialogues between the student and TeachMed. The expert model is represented by an influence diagram capturing the relationship between evidence and hypotheses related to a clinical case. RESULTS: This approach is implemented into TeachMed, a patient simulator we are developing to support clinical reasoning learning for a problem-based learning medical curriculum at our institution; we demonstrate some scenarios of tutoring feedback generated using this approach. CONCLUSION: Each of the knowledge representation formalisms that we use has already been proven successful in different applications of artificial intelligence and software engineering, but their integration into a coherent pedagogic model as we propose is unique. The examples we discuss illustrate the effectiveness of this approach, making it promising for the development of complex ITS, not only for clinical reasoning learning, but potentially for other domains as well.

Computer Simulation↗

Integrated approach for designing medical decision support systems with knowledge extracted from clinical databases by statistical methods.

In clinical research data is often studied by a particular method without previous analysis of quality or semantic contents which could link clinical database and data analytical (e.g. statistical) procedures. In order to avoid bias caused by this situation, we propose that the analysis of medical data should be divided into two main steps. In the first one we concentrate on conducting the quality, semantic and structure analyses. In the second step our aim is to build an appropriate dictionary of data analysis methods for further knowledge extraction. Methods like robust statistical techniques, procedures for mixed continuous and discrete data, fuzzy linguistic approach, machine learning and neural networks can be included. The results may be evaluated both using test samples and applying other relevant data-analytical techniques to the particular problem under the study.

Artificial Intelligence↗

Foundations of Artificial Intelligence in Hepatology: What a Clinician Needs to Know.

This review focuses on foundational knowledge about artificial intelligence (AI) in hepatology, exploring how AI, including machine learning and deep learning, leverages large-scale clinical data to transform the diagnosis, risk assessment, prognostication, and management of liver diseases. Online resources are described to offer fundamental AI knowledge and essential technical skills and to facilitate clinician participation across the entire AI lifecycle, ensuring they contribute not only as end users but also in development and deployment. Unlike traditional statistical approaches that prioritize interpretable parameters and clinical insight, AI focuses on maximizing predictive accuracy by identifying complex, often non-linear patterns using high-dimensional data, albeit often at the cost of model interpretability. AI is demonstrating clinical utility in liver histopathology and radiological imaging, significantly improving detection accuracy for cirrhosis, clinically significant portal hypertension, and hepatocellular carcinoma. Beyond diagnostics, AI-driven prediction models are emerging to provide personalized risk stratification for the development of liver-related complications and treatment guidance, based on complex data including longitudinal laboratory results, comorbidities, and co-medication use to monitor disease progression and therapy response. The field is rapidly expanding into novel areas such as analyzing patient-reported outcomes, genomic data, and real-time liver function monitoring, offering deeper mechanistic insights alongside clinical tools. Despite the potential to revolutionize hepatology practice and research, successful integration into routine care faces challenges. These include seamless workflow integration with existing electronic health records, establishing clear liability frameworks, and guaranteeing protection of patient privacy. Addressing these hurdles requires collaborative efforts from clinicians, researchers, and regulators to develop best practices and governance. Understanding the transformative capabilities, current applications, emerging frontiers, and essential implementation considerations is crucial for clinicians navigating the evolving AI landscape and responsibly utilizing its power for improved patient outcomes.

PROBAST+AI↗

Structural semantic interconnections: a knowledge-based approach to word sense disambiguation.

Word Sense Disambiguation (WSD) is traditionally considered an Al-hard problem. A break-through in this field would have a significant impact on many relevant Web-based applications, such as Web information retrieval, improved access to Web services, information extraction, etc. Early approaches to WSD, based on knowledge representation techniques, have been replaced in the past few years by more robust machine learning and statistical techniques. The results of recent comparative evaluations of WSD systems, however, show that these methods have inherent limitations. On the other hand, the increasing availability of large-scale, rich lexical knowledge resources seems to provide new challenges to knowledge-based approaches. In this paper, we present a method, called structural semantic interconnections (SSI), which creates structural specifications of the possible senses for each word in a context and selects the best hypothesis according to a grammar G, describing relations between sense specifications. Sense specifications are created from several available lexical resources that we integrated in part manually, in part with the help of automatic procedures. The SSI algorithm has been applied to different semantic disambiguation problems, like automatic ontology population, disambiguation of sentences in generic texts, disambiguation of words in glossary definitions. Evaluation experiments have been performed on specific knowledge domains (e.g., tourism, computer networks, enterprise interoperability), as well as on standard disambiguation test sets.

Algorithms↗

CoMed: a real-time collaborative medicine system.

CoMed is a prototype of a real-time collaborative medicine system that allows medical specialists to share patient records and to communicate with each other on the Internet. CoMed consists of a multimedia medical database containing relevant information about laryngeal diseases and a real-time collaboration system including a teleconferencing system, a whiteboard and a chatting system. CoMed is web-based. We adopted the object database O2 and CORBA technologies for the multimedia medical database. Therefore, our system can provide the flexibility, extensibility and location transparency of patient databases. We developed a SeeYou Active X control for the teleconferencing system and a Java applet for the whiteboard and chatting system. CoMed improves the efficiency of the overall system by separating the servers on a UNIX machine and a Windows NT machine. CoMed can be utilized for stand-alone research, for collaborative consultations among medical specialists and for a telemedicine in participation with the patients and medical specialists. Our system can be extended easily into other types of the collaborative systems, such as collaborative distance learning, collaborative science system, etc.

Humans↗

Joint learning of gene functions--a Bayesian network model approach.

In this paper, we develop a machine learning system for determining gene functions from heterogeneous data sources using a Weighted Naive Bayesian network (WNB). The knowledge of gene functions is crucial for understanding many fundamental biological mechanisms such as regulatory pathways, cell cycles and diseases. Our major goal is to accurately infer functions of putative genes or Open Reading Frames (ORFs) from existing databases using computational methods. However, this task is intrinsically difficult since the underlying biological processes represent complex interactions of multiple entities. Therefore, many functional links would be missing when only one or two sources of data are used in the prediction. Our hypothesis is that integrating evidence from multiple and complementary sources could significantly improve the prediction accuracy. In this paper, our experimental results not only suggest that the above hypothesis is valid, but also provide guidelines for using the WNB system for data collection, training and predictions. The combined training data sets contain information from gene annotations, gene expressions, clustering outputs, keyword annotations, and sequence homology from public databases. The current system is trained and tested on the genes of budding yeast Saccharomyces cerevisiae. Our WNB model can also be used to analyze the contribution of each source of information toward the prediction performance through the weight training process. The contribution analysis could potentially lead to significant scientific discovery by facilitating the interpretation and understanding of the complex relationships between biological entities.

Artificial Intelligence↗

Selection of patient samples and genes for outcome prediction.

Gene expression profiles with clinical outcome data enable monitoring of disease progression and prediction of patient survival at the molecular level. We present a new computational method for outcome prediction. Our idea is to use an informative subset of original training samples. This subset consists of only short-term survivors who died within a short period and long-term survivors who were still alive after a long follow-up time. These extreme training samples yield a clear platform to identify genes whose expression is related to survival. To find relevant genes, we combine two feature selection methods -- entropy measure and Wilcoxon rank sum test -- so that a set of sharp discriminating features are identified. The selected training samples and genes are then integrated by a support vector machine to build a prediction model, by which each validation sample is assigned a survival/relapse risk score for drawing Kaplan-Meier survival curves. We apply this method to two data sets: diffuse large-B-cell lymphoma (DLBCL) and primary lung adenocarcinoma. In both cases, patients in high and low risk groups stratified by our risk scores are clearly distinguishable. We also compare our risk scores to some clinical factors, such as International Prognostic Index score for DLBCL analysis and tumor stage information for lung adenocarcinoma. Our results indicate that gene expression profiles combined with carefully chosen learning algorithms can predict patient survival for certain diseases.

Biomarkers, Tumor↗