PubMed HealthSearch

Biomedical subjects

Li Shen

Publications and source records attributed to Li Shen.

7 recordsLinked to original sources

Sex-specific biological aging clocks across organs and omics.

Sex differentially shapes aging, neurodevelopment and neurodegenerative diseases such as Alzheimer's disease (AD). However, most biological aging clocks (artificial intelligence-predicted age minus chronological age) were trained on sex-pooled samples and implicitly assume sex invariance.Here we developed 38 sex-specific biological aging clocks across 15 organ systems. We first demonstrate the importance of sex-stratified training for constructing sex-specific healthy normative references and then reveal marked divergence between female and male clocks. Key genetic parameters and Mendelian randomization results indicate that organ-specific aging liability and its relationships to cardiometabolic, endocrine and mental traits are configured differently in females and males. Proteomic analyses identify distinct, organ-resolved synaptic, immune, vascular and metabolic networks that differentially track female and male biological aging. In longitudinal survival analyses, sex-specific clocks predict whole-body systemic diseases and all-cause mortality in a sex-dependent and organ-dependent manner. Further analyses reveal sex-dependent associations between the brain aging clock and cognitive decline trajectory during a preclinical AD clinical trial. Sex-stratified clocks may offer distinct value by defining biological age against sex-appropriate normative references and revealing sex-dependent genetic, molecular and clinical signatures that pooled models may obscure. Meanwhile, sex-pooled and sex-interaction approaches remain valuable, as human aging and disease also share fundamental biological similarities between females and males. Together, these findings reveal sex-specific biological aging signatures in aging, AD and systemic health, highlighting the need for explicitly sex-stratified modeling approaches.

Journal Article

Structural characterization and predicted biosynthetic pathway of the polysaccharide component of bioflocculant from starch-degrading Bacillus subtilis ZHX3.

Polysaccharides-based bioflocculant is a promising eco-friendly alternative to conventional flocculants, yet their application is limited by high production cost. Understanding the biosynthetic pathway is essential for targeted strain improvement. In this study, we characterized polysaccharides structure of bioflocculant MBF-ZHX3 from Bacillus subtilis ZHX3 and predicted its biosynthetic pathway via genomic analysis combined with quantitative real-time PCR (qPCR). Two purified polysaccharide fractions, PS1-1 (5982 Da) and PS2-1 (17,577 Da), were obtained. Both were mainly composed of glucose, with a backbone of →4)-α-D-Glcp-(1 → and α-D-Glcp-(1 → branches attached at O-6. Whole-genome sequencing revealed a circular chromosome of 4,122,369 bp and two plasmids. Functional annotation showed high carbohydrate metabolism activity, with 284 genes (9.52%) and 264 genes (11.28%) assigned to carbohydrate metabolism in the COG and KEGG database, respectively. A complete eps gene cluster consisting of 15 open reading frames was identified. qPCR showed that key genes involved in substrate uptake (ptsG, malP, mdxEFG-msmX) and nucleotide sugar synthesis (pgcA, gtaB) were significantly upregulated. The priming glycosyltransferase (GT) epsL and the primary GT epsF were upregulated, along with the flippase epsK, polymerase epsG, and chain-length regulators epsA and epsB. Based on these findings, we propose a putative biosynthetic pathway for the polysaccharide component of MBF-ZHX3, and identify epsL, epsF, and epsG as prioritized targets for future genetic engineering. This work provides an integrated structural-genomic-transcriptomic framework that can guide rational strain improvement to enhance bioflocculant production.

Polysaccharides structure

E2F3a transcription factor mediates behavioral, cellular, and DNA-protein regulation of cocaine reward in the nucleus accumbens.

Drug addiction is characterized by orchestrated transcriptional changes in brain reward regions, including the nucleus accumbens (NAc). The transcription factor E2F3a has emerged as a novel regulator of cocaine's rewarding effects, yet its sex- and cell-specific mechanisms, as well as its genome-wide targets, remain undetermined. Here, we investigated the motivational and reinforcing roles of E2F3a in cocaine reward using conditioned place preference (CPP) and self-administration, combined with behavioral economics and viral-mediated gene manipulation. Selective overexpression of E2F3a in D1-type medium spiny neurons (MSNs), but not D2-MSNs, increased cocaine CPP in both male and female mice, whereas knockdown produced the opposite effects. Behavioral economics analyses further revealed that E2F3a regulates specific aspects of cocaine reinforcement. Genome-wide mapping revealed increased E2F3a binding to DNA at genes associated with cocaine exposure. Together, these results establish E2F3a as a central substrate of cocaine reward via the recruitment of D1-MSNs and coordinated expression of both proven and new molecular drivers.

Journal Article

Cross-Device Adaptation of Mirai for Mammography-Based Breast Cancer Risk Prediction.

Fine-tuning can adapt pretrained medical imaging models to new clinical datasets, but device-specific domain shifts may limit generalizability. We evaluated Mirai, a mammography-based deep learning model for breast cancer risk prediction, in a large screening cohort containing Hologic and General Electric (GE) full-field digital mammography systems, including GE Premium View (GE PV) and Tissue Equalization (GE TE) post-processing software. Native Mirai showed lower performance on TE images than on Hologic or PV images. Fine-tuning on TE images improved TE performance, particularly for short-term risk prediction, but substantially reduced performance on Hologic images, consistent with catastrophic forgetting. To mitigate this effect, we developed a device-invariant model using interleaved multi-device sampling and conditional adversarial training. This approach largely restored Hologic performance while maintaining improved TE performance, providing better robustness across heterogeneous imaging platforms. Comparison of cumulative and annual risk AUCs over a five-year time horizon further showed that performance gains were driven mainly by short- and intermediate-term predictions. These findings highlight both the value and dangers of device-specific fine-tuning and support balanced domain-adaptation strategies for deploying mammography-based risk models across diverse clinical imaging environments.

Journal Article

JASMINE: A powerful representation learning method for enhanced analysis of incomplete multi-omics data.

Integrative analysis of multi-omics data provides a more comprehensive and nuanced view of a subject's biological state. However, high-dimensionality and ubiquitous modality missingness present significant analytical challenges. Existing methods for incomplete multi-omics data are scarce, do not fully leverage both modality-specific and shared information, and produce task-biased representations. We propose JASMINE, a self-supervised representation learning method for incomplete multi-omics data that preserves both modality-specific and joint information and enhances sample similarity structure. JASMINE produces embeddings that achieve superior performance across multiple tasks for two different incomplete multi-omics datasets while requiring only a single round of training per dataset.

missing data

AlphaGenome Enhances Personal Gene Expression Prediction but Retains Key Limitations.

In recent years, numerous genome AI models have been developed to elucidate the relationship between DNA sequence and gene expression. However, these models have faced criticism for their limited accuracy in predicting individual-specific gene expression. AlphaGenome, the current state-of-the-art in genome AI, achieves exceptional performance across a range of sequence-based predictive tasks, but its utility for personal expression prediction has not yet been assessed. In this study, we evaluate AlphaGenome's ability to predict personal gene expression and find that it significantly outperforms its predecessor. Using GTEx data, AlphaGenome improves the prediction of expression direction over Enformer, achieving an odds ratio of 3.0. In some cases, it even reverses previously observed negative correlations into positive ones. Moreover, AlphaGenome demonstrates improved performance for genes with known nonlinear sequence-expression relationships, though it uncovers mechanisms distinct from those identified by tree-based models.

deep learning

S-GMAS: Genome-Wide Mediation Analysis With Brain Subcortical Shape Mediators.

Mediation analysis is widely utilized in neuroscience to investigate the role of brain image phenotypes in the neurological pathways from genetic exposures to clinical outcomes. However, it is still difficult to conduct mediation analyses with whole genome-wide exposures and brain subcortical shape mediators due to several challenges including (i) large-scale genetic exposures, that is, millions of single-nucleotide polymorphisms (SNPs); (ii) nonlinear Hilbert space for shape mediators; and (iii) statistical inference on the direct and indirect effects. To tackle these challenges, this paper proposes a genome-wide mediation analysis framework with brain subcortical shape mediators. First, to address the issue caused by the high dimensionality in genetic exposures, a fast genome-wide association analysis is conducted to discover potential genetic variants with significant genetic effects on the clinical outcome. Second, the square-root velocity function representations are extracted from the brain subcortical shapes, which fall in an unconstrained linear Hilbert subspace. Third, to identify the underlying causal pathways from the detected SNPs to the clinical outcome implicitly through the shape mediators, we utilize a shape mediation analysis framework consisting of a shape-on-scalar model and a scalar-on-shape model. Furthermore, the bootstrap resampling approach is adopted to investigate both global and spatial significant mediation effects. Finally, our framework is applied to the corpus callosum shape data from the Alzheimer's Disease Neuroimaging Initiative.

Humans