PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “penalized regression”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

PROLONG: penalized regression for outcome guided longitudinal omics analysis with network and group constraints.

MOTIVATION: There is a growing interest in longitudinal omics data paired with some longitudinal clinical outcome. Given a large set of continuous omics variables and some continuous clinical outcome, each measured for a few subjects at only a few time points, we seek to identify those variables that co-vary over time with the outcome. To motivate this problem we study a dataset with hundreds of urinary metabolites along with Tuberculosis mycobacterial load as our clinical outcome, with the objective of identifying potential biomarkers for disease progression. For such data clinicians usually apply simple linear mixed effects models which often lack power given the low number of replicates and time points. We propose a penalized regression approach on the first differences of the data that extends the lasso + Laplacian method [Li and Li (Network-constrained regularization and variable selection for analysis of genomic data. Bioinformatics 2008;24:1175-82.)] to a longitudinal group lasso + Laplacian approach. Our method, PROLONG, leverages the first differences of the data to increase power by pairing the consecutive time points. The Laplacian penalty incorporates the dependence structure of the variables, and the group lasso penalty induces sparsity while grouping together all contemporaneous and lag terms for each omic variable in the model. RESULTS: With an automated selection of model hyper-parameters, PROLONG correctly selects target metabolites with high specificity and sensitivity across a wide range of scenarios. PROLONG selects a set of metabolites from the real data that includes interesting targets identified during EDA. AVAILABILITY AND IMPLEMENTATION: An R package implementing described methods called "prolong" is available at https://github.com/stevebroll/prolong. Code snapshot available at 10.5281/zenodo.14804245.

Humans↗

Individual and population penalized regression splines for accelerated longitudinal designs.

In an accelerated longitudinal design (ALD), individuals enter the study at different points of their growth trajectory and are observed over a short time span relative to the entire time span of interest. ALD data are combined across independent units to provide an estimate of an overall population curve and predictions of individual patterns of change. As a modest extension of the work of Ruppert et al. (2003, Semiparametric Regression, Cambridge University Press), we develop a computationally efficient procedure for the application of longitudinal semiparametric methods under ALD sampling schemes. We compare balanced and complete longitudinal designs to ALDs using the Berkeley Growth Study data and apply our method to longitudinal magnetic resonance imaging (MRI) brain structure size (volume) measurements from an ongoing developmental study. Potential applications extend beyond growth studies to many other fields in which cost and feasibility constraints impose restrictions on sample size and on the numbers and timings of repeated measurements across subjects.

Adolescent↗

Bayesian proportional hazards model with time-varying regression coefficients: a penalized Poisson regression approach.

One can fruitfully approach survival problems without covariates in an actuarial way. In narrow time bins, the number of people at risk is counted together with the number of events. The relationship between time and probability of an event can then be estimated with a parametric or semi-parametric model. The number of events observed in each bin is described using a Poisson distribution with the log mean specified using a flexible penalized B-splines model with a large number of equidistant knots. Regression on pertinent covariates can easily be performed using the same log-linear model, leading to the classical proportional hazard model. We propose to extend that model by allowing the regression coefficients to vary in a smooth way with time. Penalized B-splines models will be proposed for each of these coefficients. We show how the regression parameters and the penalty weights can be estimated efficiently using Bayesian inference tools based on the Metropolis-adjusted Langevin algorithm.

Bayes Theorem↗

Differential gene expression detection and sample classification using penalized linear regression models.

Differential gene expression detection and sample classification using microarray data have received much research interest recently. Owing to the large number of genes p and small number of samples n (p >> n), microarray data analysis poses big challenges for statistical analysis. An obvious problem owing to the 'large p small n' is over-fitting. Just by chance, we are likely to find some non-differentially expressed genes that can classify the samples very well. The idea of shrinkage is to regularize the model parameters to reduce the effects of noise and produce reliable inferences. Shrinkage has been successfully applied in the microarray data analysis. The SAM statistics proposed by Tusher et al. and the 'nearest shrunken centroid' proposed by Tibshirani et al. are ad hoc shrinkage methods. Both methods are simple, intuitive and prove to be useful in empirical studies. Recently Wu proposed the penalized t/F-statistics with shrinkage by formally using the (1) penalized linear regression models for two-class microarray data, showing good performance. In this paper we systematically discussed the use of penalized regression models for analyzing microarray data. We generalize the two-class penalized t/F-statistics proposed by Wu to multi-class microarray data. We formally derive the ad hoc shrunken centroid used by Tibshirani et al. using the (1) penalized regression models. And we show that the penalized linear regression models provide a rigorous and unified statistical framework for sample classification and differential gene expression detection.

Algorithms↗

Dimension reduction-based penalized logistic regression for cancer classification using microarray data.

The use of penalized logistic regression for cancer classification using microarray expression data is presented. Two dimension reduction methods are respectively combined with the penalized logistic regression so that both the classification accuracy and computational speed are enhanced. Two other machine-learning methods, support vector machines and least-squares regression, have been chosen for comparison. It is shown that our methods have achieved at least equal or better results. They also have the advantage that the output probability can be explicitly given and the regression coefficients are easier to interpret. Several other aspects, such as the selection of penalty parameters and components, pertinent to the application of our methods for cancer classification are also discussed.

Algorithms↗

Differential gene expression detection using penalized linear regression models: the improved SAM statistics.

UNLABELLED: Differential gene expression detection using microarrays has received lots of research interests recently. Many methods have been proposed, including variants of F-statistics, non-parametric approaches and empirical Bayesian methods etc. The SAM statistics has been shown to have good performance in empirical studies. SAM is more like an ad hoc shrinkage method. The idea is that for small sample microarray data, it is often useful to pool information across genes to improve efficiency. Under Bayesian framework Smyth formally derived the test statistics with shrinkage using the hierarchical models. In this paper we cast differential gene expression detection in the familiar framework of linear regression model. Commonly used test statistics correspond to using least squares to estimate the regression parameters. Based on the vast literature of research on linear models, we can naturally consider other alternatives. Here we explore the penalized linear regression. We propose the penalized t-/F-statistics for two-class microarray data based on [Formula: see text] penalty. We will show that the penalized test statistics intuitively makes sense and through applications we illustrate its good performance. AVAILABILITY: Supplementary information including program codes, more detailed analysis results and R functions for the proposed methods can be found at http://www.biostat.umn.edu/~baolin/research CONTACT: baolin@biostat.umn.edu SUPPLEMENTARY INFORMATION: http://www.biostat.umn.edu/~baolin/research.

Cell Line, Tumor↗

Penalized binary regression for gene expression profiling.

OBJECTIVES: A typical bioinformatics task in microarray analysis is the classification of biological samples into two alternative categories. A procedure is needed which, based on the expression levels measured, allows us to compute the probability that a new sample belongs to a certain class. METHODS: For the purpose of classification the statistical approach of binary regression is considered. Highdimensionality and at the same time small sample sizes make it a challenging task. Standard logit or probit regression fails because of condition problems and poor predictive performance. The concepts of frequentist and of Bayesian penalization for binary regression are introduced. A Bayesian interpretation of the penalized log-likelihood is given. Finally the role of cross-validation for regularization and feature selection is discussed. RESULTS: Penalization makes classical binary regression a suitable tool for microarray analysis. We illustrate penalized logit and Bayesian probit regression on a well-known data set and compare the obtained results, also with respect to published results from decision trees. CONCLUSIONS: The frequentist and the Bayesian penalization concept work equally well on the example data, however some method-specific differences can be made out. Moreover the Bayesian approach yields a quantification (posterior probabilities) of the bias due to the constraining assumptions.

Bayes Theorem↗

Classification using partial least squares with penalized logistic regression.

MOTIVATION: One important aspect of data-mining of microarray data is to discover the molecular variation among cancers. In microarray studies, the number n of samples is relatively small compared to the number p of genes per sample (usually in thousands). It is known that standard statistical methods in classification are efficient (i.e. in the present case, yield successful classifiers) particularly when n is (far) larger than p. This naturally calls for the use of a dimension reduction procedure together with the classification one. RESULTS: In this paper, the question of classification in such a high-dimensional setting is addressed. We view the classification problem as a regression one with few observations and many predictor variables. We propose a new method combining partial least squares (PLS) and Ridge penalized logistic regression. We review the existing methods based on PLS and/or penalized likelihood techniques, outline their interest in some cases and theoretically explain their sometimes poor behavior. Our procedure is compared with these other classifiers. The predictive performance of the resulting classification rule is illustrated on three data sets: Leukemia, Colon and Prostate.

Algorithms↗

Penalized Cox regression analysis in the high-dimensional and low-sample size settings, with applications to microarray gene expression data.

MOTIVATION: An important application of microarray technology is to relate gene expression profiles to various clinical phenotypes of patients. Success has been demonstrated in molecular classification of cancer in which the gene expression data serve as predictors and different types of cancer serve as a categorical outcome variable. However, there has been less research in linking gene expression profiles to the censored survival data such as patients' overall survival time or time to cancer relapse. It would be desirable to have models with good prediction accuracy and parsimony property. RESULTS: We propose to use the L(1) penalized estimation for the Cox model to select genes that are relevant to patients' survival and to build a predictive model for future prediction. The computational difficulty associated with the estimation in the high-dimensional and low-sample size settings can be efficiently solved by using the recently developed least-angle regression (LARS) method. Our simulation studies and application to real datasets on predicting survival after chemotherapy for patients with diffuse large B-cell lymphoma demonstrate that the proposed procedure, which we call the LARS-Cox procedure, can be used for identifying important genes that are related to time to death due to cancer and for building a parsimonious model for predicting the survival of future patients. The LARS-Cox regression gives better predictive performance than the L(2) penalized regression and a few other dimension-reduction based methods. CONCLUSIONS: We conclude that the proposed LARS-Cox procedure can be very useful in identifying genes relevant to survival phenotypes and in building a parsimonious predictive model that can be used for classifying future patients into clinically relevant high- and low-risk groups based on the gene expression profile and survival times of previous patients.

Biomarkers, Tumor↗

Threshold gradient descent method for censored data regression with applications in pharmacogenomics.

An important area of research in pharmacogenomics is to relate high-dimensional genetic or genomic data to various clinical phenotypes of patients. Due to large variability in time to certain clinical event among patients, studying possibly censored survival phenotypes can be more informative than treating the phenotypes as categorical variables. In this paper, we develop a threshold gradient descent (TGD) method for the Cox model to select genes that are relevant to patients' survival and to build a predictive model for the risk of a future patient. The computational difficulty associated with the estimation in the high-dimensional and low-sample size settings can be efficiently solved by the gradient descent iterations. Results from application to real data set on predicting survival after chemotherapy for patients with diffuse large B-cell lymphoma demonstrate that the proposed method can be used for identifying important genes that are related to time to death due to cancer and for building a parsimonious model for predicting the survival of future patients. The TGD based Cox regression gives better predictive performance than the L2 penalized regression and can select more relevant genes than the L1 penalized regression.

Algorithms↗

Classification of gene microarrays by penalized logistic regression.

Classification of patient samples is an important aspect of cancer diagnosis and treatment. The support vector machine (SVM) has been successfully applied to microarray cancer diagnosis problems. However, one weakness of the SVM is that given a tumor sample, it only predicts a cancer class label but does not provide any estimate of the underlying probability. We propose penalized logistic regression (PLR) as an alternative to the SVM for the microarray cancer diagnosis problem. We show that when using the same set of genes, PLR and the SVM perform similarly in cancer classification, but PLR has the advantage of additionally providing an estimate of the underlying probability. Often a primary goal in microarray cancer diagnosis is to identify the genes responsible for the classification, rather than class prediction. We consider two gene selection methods in this paper, univariate ranking (UR) and recursive feature elimination (RFE). Empirical results indicate that PLR combined with RFE tends to select fewer genes than other methods and also performs well in both cross-validation and test samples. A fast algorithm for solving PLR is also described.

Algorithms↗

Growth velocity assessment in paediatric AIDS: smoothing, penalized quantile regression and the definition of growth failure.

The analysis of growth records in paediatric anti-HIV clinical trials plays an important role in trial evaluation. Growth failure may be a manifestation of progressive disease or treatment toxicity, and is commonly specified as a major trial outcome event indicating poor treatment performance. Despite new therapeutic advances against HIV proliferation in infected patients, accurate monitoring and interpretation of somatic growth in paediatric AIDS remains clinically important, in light of uncertainties regarding relationship between viral load reductions and achievement of favourable somatic growth profiles. Our aim in this paper is to construct a criterion for growth failure that discriminates patients whose risk of death subsequent to growth failure is elevated to a clinically significant degree. To construct the criterion, individual growth curves and velocities are modelled using loess smoothing, penalized likelihood quantile regressions are fit to model age-specific growth velocity distributions for gender-stratified cohorts, and proportional hazards model selection is conducted to identify features of velocity series that are informative on the survival distribution. The resulting growth-failure criterion is expected to be useful for disease staging in resource-limited medical environments where T-cell counts and viral load measures are unavailable.

Adolescent↗

Generalized additive modeling with implicit variable selection by likelihood-based boosting.

The use of generalized additive models in statistical data analysis suffers from the restriction to few explanatory variables and the problems of selection of smoothing parameters. Generalized additive model boosting circumvents these problems by means of stagewise fitting of weak learners. A fitting procedure is derived which works for all simple exponential family distributions, including binomial, Poisson, and normal response variables. The procedure combines the selection of variables and the determination of the appropriate amount of smoothing. Penalized regression splines and the newly introduced penalized stumps are considered as weak learners. Estimates of standard deviations and stopping criteria, which are notorious problems in iterative procedures, are based on an approximate hat matrix. The method is shown to be a strong competitor to common procedures for the fitting of generalized additive models. In particular, in high-dimensional settings with many nuisance predictor variables it performs very well.

Biometry↗

Semiparametric models for missing covariate and response data in regression models.

We consider a class of semiparametric models for the covariate distribution and missing data mechanism for missing covariate and/or response data for general classes of regression models including generalized linear models and generalized linear mixed models. Ignorable and nonignorable missing covariate and/or response data are considered. The proposed semiparametric model can be viewed as a sensitivity analysis for model misspecification of the missing covariate distribution and/or missing data mechanism. The semiparametric model consists of a generalized additive model (GAM) for the covariate distribution and/or missing data mechanism. Penalized regression splines are used to express the GAMs as a generalized linear mixed effects model, in which the variance of the corresponding random effects provides an intuitive index for choosing between the semiparametric and parametric model. Maximum likelihood estimates are then obtained via the EM algorithm. Simulations are given to demonstrate the methodology, and a real data set from a melanoma cancer clinical trial is analyzed using the proposed methods.

Algorithms↗

Pretreatment EBV-DNA/TLG-Based Risk Stratification Is Associated With Survival Outcomes in Nonmetastatic Nasopharyngeal Carcinoma: An Exploratory Study.

Whether combining pretreatment plasma Epstein-Barr virus DNA (EBV-DNA) with 18F-FDG PET/CT-derived total lesion glycolysis (TLG) improves prognostic stratification in nonmetastatic nasopharyngeal carcinoma (NPC) is unclear, particularly in nonendemic populations. We retrospectively analyzed 86 eligible nonmetastatic NPC patients treated with definitive radiotherapy (2010-2024) at a single nonendemic-region institution. EBV-DNA (prespecified cutoff 3500 copies/mL) and TLG (cutoff 200, ROC-derived within this cohort) were dichotomized. Both were available in 59/86 patients (68.6%), who differed from the rest in nodal and overall stage and in RT technique. Baseline PET/CT was in-house in 57 of 86 patients, and a robustness analysis in that subgroup is reported. Given limited events (13 PFS, 9 OS), Cox analyses are exploratory and were supplemented with penalized regression and bootstrap validation. At a median follow-up of 75.5 months, 5-year PFS and OS for the whole cohort (n = 86) were 81.1% and 85.9%. The EBV-DNAhigh/TLGhigh subgroup remained associated with inferior PFS after adjustment in an exploratory model (adjusted HR = 3.97, 95% CI: 1.32-11.93) and, in a single-variable model, with inferior OS (HR = 4.13, 95% CI: 1.10-15.52). Discrimination was comparable to the individual-biomarker model for PFS and lower for OS. Five-year PFS fell monotonically across the four risk groups in the complete-case cohort (n = 59; 89.7%-58.3%). OS differed across groups (log-rank p = 0.044) but was not strictly monotonic, with wide, overlapping confidence intervals. This two-biomarker model is hypothesis-generating and needs prospective, multicenter validation before any consideration of risk-adapted treatment.

Epstein–Barr virus DNA↗

Identification of a novel human gut microbes and microbial metabolites related genes signature for prognostic implication in head and neck squamous carcinomas.

BACKGROUND: The gut microbiota acts as a critical driver influencing the pathogenesis, therapeutic response, and clinical outcomes across various cancer types. This study aimed to investigate the prognostic value of human gut microbes and microbial metabolites related genes (HGMMMRGs) in head and neck squamous cell carcinoma (HNSCC). METHODS: We constructed a prognostic risk model comprising 19 core HGMMMRGs using LASSO penalized regression and a multivariate Cox proportional hazards model. The predictive performance of the model was evaluated through Kaplan-Meier analysis, receiver operating characteristic (ROC) curves, nomograms, and concordance index. In addition, functional enrichment analysis was performed on the differentially expressed risk genes. Furthermore, the relationship between the immune microenvironment of HNSCC and the risk diagnostic model was examined. Western blot analysis was used to assess the expression levels of IL10 in both HNSCC tissues and adjacent normal tissues. Finally, the correlation between IL10 and the gut microbiota was explored. RESULTS: This study developed a risk score model integrating 19 HGMMMRG genes, which can serve as a tool to guide prognosis and immune microenvironment assessment in HNSCC patients. Survival analysis showed that patients in the high-risk group had significantly worse outcomes (P&#x2009;<&#x2009;0.05). Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis revealed significant enrichment of differentially expressed genes (DRLs) and immune-related pathways. Western blot analysis further confirmed that IL10 was highly expressed in HNSCC, and the abundance of Faecalibacterium prausnitzii and Enterococcus durans colonies was correlated with IL10 expression. CONCLUSION: We developed a prognostic model for HGMMMRGs that can be effectively used to predict OS in patients with HNSCC. Second, Faecalibacterium prausnitzii and Enterococcus durans can influence the prognosis of patients with HNSCC by mediating the expression IL10 and thereby affecting the prognosis of HNSCC patients. Thus, human gut microbes and microbial metabolite-related genes may be another promising strategy for the treatment of patients with HNSCC.

HNSCC↗

Multi-Ancestry Genome-Wide Association with Fine-Mapping Identifies Novel Loci for Pigment Dispersion Syndrome and Pigmentary Glaucoma.

PURPOSE: Pigment dispersion syndrome and pigmentary glaucoma are important causes of ocular hypertension and glaucomatous optic neuropathy, yet their genetic determinants remain incompletely defined, particularly across diverse ancestries. This study aimed to use a large multi-ancestry cohort from the All of Us Research Program to investigate the genetic basis of pigment dispersion syndrome and pigmentary glaucoma. DESIGN: Case-control study. PARTICIPANTS: In total, 572 cases and 37&#x2009;808 controls with array genotyping and 537 cases and 35&#x2009;493 controls with whole-genome sequencing. METHODS: Using electronic health record phenotyping in the All of Us Research Program, we performed multi-ancestry genome-wide association analyses using both array-based data and whole-genome sequencing-based data, comparing patients with pigment dispersion syndrome or pigmentary glaucoma to those without either condition. We also performed Firth penalized regression and Fisher analyses, and we performed principal component analyses to assess effect sizes across genetic ancestries. We applied statistical fine-mapping, examined for cross-trait overlap, and assessed expression quantitative trait locus associations for lead variants. MAIN OUTCOME MEASURES: P values and odds ratios of lead loci from genome-wide association analyses; size of credible sets determined from fine-mapping; allele frequency of lead variants in cases, controls, and the general population; expression quantitative trait loci effect size and P values linking lead variants to gene expression. RESULTS: We identified 4 loci reaching genome-wide significance across analyses, including signals near EPHA7 (which mediates cell-cell signaling), within TYR (involved in melanin synthesis and replicated from prior studies), within LINC01138, and near OTX2. Statistical fine-mapping refined 3 of these loci to single-variant 95% credible sets and narrowed the TYR locus to small credible sets, prioritizing possible causal variants. Effect estimates were broadly consistent across genetic ancestry clusters. Lead variants showed regulatory evidence in expression quantitative trait locus, including reduced EPHA7 expression. CONCLUSIONS: These findings implicate both melanogenesis and cell-cell adhesion and signaling pathways in pigment dispersion syndrome and pigmentary glaucoma. FINANCIAL DISCLOSURE(S): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Genome-wide association study↗

Sparse polygenic risk score inference with the spike-and-slab LASSO.

MOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs.

Multifactorial Inheritance↗