PubMed HealthSearch

Biomedical subjects

Joseph Glessner

Publications and source records attributed to Joseph Glessner.

3 recordsLinked to original sources

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis

Computational strategies for copy number variation detection, disease association, and beyond.

Copy number variations (CNVs) are key structural variations that contribute to human genetic diversity, evolution, and disease susceptibility. Advances in sequencing technologies and computational methods have improved CNV detection, yet association studies remain challenged by methodological limitations and a lack of standardisation. This review provides an overview of computational strategies for germline CNV detection and disease association. We highlight the value of CNV analysis for uncovering genetic contributions to complex traits and disease risk and outline an analysis workflow including key benchmarking methods. We also discuss current challenges and future directions for advancing CNV detection and association analysis.

Humans

Deep learning and statistical methods identify novel asthma risk variants in Europeans.

BACKGROUND: Asthma is a common heritable respiratory disorder with a complex genetic basis. Although large-scale genome-wide association studies have identified many risk loci, the full spectrum of its polygenic architecture remains to be defined. OBJECTIVE: We refined the genetic landscape of asthma in individuals of European ancestry and improve polygenic risk prediction through statistical and deep learning-based methods. METHODS: We conducted the largest genome-wide association study meta-analysis of asthma in individuals of European ancestry, combining data from the Global Biobank Meta-analysis Initiative (121,940 cases, 1,254,131 controls) and the Million Veteran Program (36,823 cases, 398,278 controls). To enhance discovery, we applied pleiotropy-informed multitrait analysis and conditional false discovery rate approaches, each incorporating eosinophil counts as a secondary trait. In parallel, we used a Transformer-based deep learning framework to further prioritize variants and improve polygenic risk prediction. RESULTS: The meta-analysis identified 69 independent genome-wide significant loci (P&#x2009;<&#x2009;5 &#xd7; 10-8) not previously reported in asthma. Multitrait analysis of genome-wide association studies, conditional false discovery rate, and deep learning approaches uncovered additional candidate loci. Functional annotation and expression quantitative trait locus mapping implicated novel genes in immune regulation, airway remodeling, and metabolic processes. Polygenic risk score models derived from deep learning-prioritized variants outperformed those based on conventional genome-wide association study and standard statistical approaches. CONCLUSIONS: Our study yields a comprehensive map of asthma-associated loci in European ancestry populations, improves genetic risk prediction, and informs future mechanistic studies.

Humans