PubMed HealthSearch

SEARCH · PubMed Health

Results for “De novo design”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

De novo designed Hsp70 activator dissolves intracellular condensates.

Protein quality control (PQC) is carried out in part by the chaperone Hsp70 in concert with adapters of the J-domain protein (JDP) family. The JDPs, also called Hsp40s, are thought to recruit Hsp70 into complexes with specific client proteins. However, the molecular principles regulating this process are not well understood. We describe the de novo design of Hsp70 binding proteins that either inhibit or stimulate Hsp70 ATPase activity. An ATPase stimulating design promoted the refolding of denatured luciferase in vitro, similar to native JDPs. Targeting of this design to intracellular condensates resulted in their nearly complete dissolution and revealed roles as cell growth promoting signaling hubs. The designs inform our understanding of chaperone structure-function relationships and provide a general and modular way to target PQC systems to regulate condensates and other cellular targets.

HSP70 Heat-Shock Proteins

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design

Predictive design of tissue-specific mammalian enhancers that function in the mouse embryo.

Enhancers control tissue-specific gene expression across animals1. Although deep learning2,3 has enabled enhancer prediction and design in mammalian cell lines and non-mammalian model organisms4-10 (reviewed in a previous publication11), it remains unclear whether such approaches can operate within the regulatory complexity of mammalian genomes and tissues in vivo. Here we present a general strategy for designing tissue-specific enhancers that function reliably in mice. We use deep learning to train compact convolutional neural networks on curated chromatin accessibility data and fine-tune them by transfer learning on validated human and mouse enhancers. Guided by these models, we design 15 synthetic enhancers for the heart, limb and central nervous system in mouse embryos, all of which are active in their intended target tissue. These results demonstrate that mammalian enhancer function can be reliably inferred from DNA sequence alone, enabling the predictive de novo design of tissue-specific synthetic enhancers from modest training sets. This work establishes a generalizable framework for programmable control of mammalian gene expression in vivo, opening new avenues in functional genomics, synthetic biology and gene therapy.

Animals

Genome-wide computational analysis reveals cardiomyocyte-specific transcriptional Cis-regulatory motifs that enable efficient cardiac gene therapy.

Gene therapy is a promising emerging therapeutic modality for the treatment of cardiovascular diseases and hereditary diseases that afflict the heart. Hence, there is a need to develop robust cardiac-specific expression modules that allow for stable expression of the gene of interest in cardiomyocytes. We therefore explored a new approach based on a genome-wide bioinformatics strategy that revealed novel cardiac-specific cis-acting regulatory modules (CS-CRMs). These transcriptional modules contained evolutionary-conserved clusters of putative transcription factor binding sites that correspond to a "molecular signature" associated with robust gene expression in the heart. We then validated these CS-CRMs in vivo using an adeno-associated viral vector serotype 9 that drives a reporter gene from a quintessential cardiac-specific α-myosin heavy chain promoter. Most de novo designed CS-CRMs resulted in a >10-fold increase in cardiac gene expression. The most robust CRMs enhanced cardiac-specific transcription 70- to 100-fold. Expression was sustained and restricted to cardiomyocytes. We then combined the most potent CS-CRM4 with a synthetic heart and muscle-specific promoter (SPc5-12) and obtained a significant 20-fold increase in cardiac gene expression compared to the cytomegalovirus promoter. This study underscores the potential of rational vector design to improve the robustness of cardiac gene therapy.

Animals

NovoBoard: A Comprehensive Framework for Evaluating the False Discovery Rate and Accuracy of De Novo Peptide Sequencing.

De novo peptide sequencing is one of the most fundamental research areas in mass spectrometry-based proteomics. Many methods have often been evaluated using a couple of simple metrics that do not fully reflect their overall performance. Moreover, there has not been an established method to estimate the false discovery rate (FDR) of de novo peptide-spectrum matches. Here we propose NovoBoard, a comprehensive framework to evaluate the performance of de novo peptide-sequencing methods. The framework consists of diverse benchmark datasets (including tryptic, nontryptic, immunopeptidomics, and different species) and a standard set of accuracy metrics to evaluate the fragment ions, amino acids, and peptides of the de novo results. More importantly, a new approach is designed to evaluate de novo peptide-sequencing methods on target-decoy spectra and to estimate and validate their FDRs. Our FDR estimation provides valuable information to assess the reliability of new peptides identified by de novo sequencing tools, especially when no ground-truth information is available to evaluate their accuracy. The FDR estimation can also be used to evaluate the capability of de novo peptide sequencing tools to distinguish between de novo peptide-spectrum matches and random matches. Our results thoroughly reveal the strengths and weaknesses of different de novo peptide-sequencing methods and how their performances depend on specific applications and the types of data.

Peptides

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis

EPIC: multi-objective guided diffusion for epitope design in TCR-pMHC complexes.

MOTIVATION: T cell receptor (TCR) recognition of peptide-major histocompatibility complex (pMHC) complexes is central to adaptive immunity, yet rational design of immunogenic epitopes remains elusive due to complex triplet binding constraints and data scarcity. No existing method can generate epitopes satisfying simultaneous requirements for antigenicity, MHC presentation, and TCR specificity. RESULTS: We present EPIC, a multi-objective diffusion framework that decomposes TCR-pMHC binding into three biologically grounded sub-tasks, enabling training-free gradient guidance without end-to-end retraining. By integrating ESM-based classifiers with a peptide diffusion generator, EPIC leverages heterogeneous immunological interaction datasets to generate diverse, context-aware epitopes. EPIC-designed top-three epitopes achieve lower predicted interface energies compared to ground-truth epitopes in 78.31% of test cases, while maintaining 80.1% sequence novelty and comparable structural confidence. Generated epitopes exhibit 100% uniqueness, high diversity (64.05%), and high antigenicity scores (0.4723). To our knowledge, EPIC is the first computational framework capable of de novo epitope design while explicitly integrating the triplet constraints of TCR-pMHC binding. This paradigm shift from discovery to design unlocks new potential for personalized cancer vaccines, precision adoptive T cell therapy, and rapid response to emerging infectious diseases. AVAILABILITY AND IMPLEMENTATION: The source code of EPIC is available at https://github.com/Octopus125/EPIC and archived on Zenodo (DOI: 10.5281/zenodo.18537646).

Receptors, Antigen, T-Cell

Application of emerging technologies in the antiviral field.

Viral diseases pose a serious threat to global public health, agriculture, and biosecurity. Conventional antiviral strategies are often limited by an incomplete understanding of disease mechanisms, poor targeting precision, and slow response times. Emerging technologies are now reshaping the landscape of antiviral research. This review examines the roles of four key frontiers, including organoid models, gene editing, AI-driven molecular design, and synthetic biology. Organoids provide physiologically relevant platforms that model virus-host interactions and disease progression. Viral infections remain a major challenge to human and animal health, agriculture, and biosecurity. Progress in antiviral research is constrained by the complexity of viral pathogenesis, the diversity and rapid evolution of viruses, and the limited translational relevance of some traditional model systems. Recent advances in organoid technology, gene editing, artificial intelligence, and synthetic biology are expanding the toolkit available for antiviral research and development. In this review, we discuss how these four technological frontiers contribute to disease modeling, target discovery, molecular design, and translational innovation. Organoids, in particular, provide physiologically relevant systems for investigating viral infection, tissue tropism, host responses, and pathogenesis. Gene editing tools, such as CRISPR, enable precise manipulation of host and viral genomes, facilitating the development of resistant organisms and next-generation vaccine platforms. AI technologies, including AlphaFold for structure prediction and platforms for de novo protein design, address long-standing bottlenecks in structural biology and offer powerful means to engineer antiviral proteins, antibodies, and vaccine antigens. Synthetic biology, guided by the Design-Build-Test-Learn cycle, integrates computational design, genetic assembly, and functional validation into a cohesive pipeline. Together, these technologies form a synergistic workflow that spans disease modeling, target discovery, molecular design, construction, testing, and iterative optimization. This integrated approach is shifting antiviral development from traditional empirical methods toward more precise, intelligent strategies. The review also highlights ongoing challenges in integration and scalability, stressing that high-quality biological datasets and stronger interdisciplinary collaboration are essential for realizing translational potential. By presenting a cohesive view of these converging methodologies, this review offers a framework to guide the intelligent evolution of antiviral strategies in both human and animal health.

Antiviral

seq2ribo: structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context. AVAILABILITY: seq2ribo is available at https://github.com/Kingsford-Group/seq2ribo.

Machine Learning

Artificial Intelligence for Natural Products Discovery and Development.

Natural products (NPs) remain a cornerstone of modern drug discovery, offering stereochemical complexity and diverse bioactivities that precisely modulate therapeutic targets, refined through billions of years of evolution. However, their research has long been hindered by inefficient, empirical workflows, high resource consumption, structural complexity, and the "multicomponent, multi-target" nature of their mechanisms. The exponential growth of genomic, metabolomic, and spectral data has overwhelmed conventional analytical methods, exposing critical bottlenecks in handling high-dimensional, heterogeneous datasets that exceed human interpretive capacity. Artificial intelligence (AI) is emerging as a transformative paradigm to address these challenges, integrating multi-omics and chemical data to shift NP research from fragmented empiricism toward mechanism-driven, precision-oriented development. By leveraging deep learning architectures- including graph neural networks, Transformers, and diffusion-based generative models-AI enables systematic decoding of NP biosynthesis, automated structure elucidation, rational target identification, knowledge extraction from vast unstructured scientific literature, and de novo molecular design. This review comprehensively surveys recent advances in AI applications across the full NP discovery and development pipeline, encompassing genome mining, structure-based and ligand-based virtual screening, multimodal structural characterization, lead optimization, and biosynthetic pathway engineering. We further examine the emerging roles of protein-centric, molecule- centric, and multimodal foundation models, as well as large language models, in bridging genotype-to-chemotype gaps and unlocking unstructured scientific knowledge. Finally, we discuss critical challenges including data scarcity, representational limitations for complex stereochemistry, physical plausibility in generative models, and the urgent need for experimental validation, while outlining future directions toward autonomous experimentation, closed-loop optimization, and human-AI collaborative discovery.

Artificial intelligence

seq2ribo: Structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context.

Journal Article

Upcycling Vegetable Waste Into Functional Food Ingredients via Synergistic Microbial Engineering and Artificial Intelligence.

The escalating generation of global vegetable waste represents a critical loss of bioactive resources, necessitating a paradigm shift from passive disposal to active nutrient upcycling. However, the industrial conversion of this heterogeneous biomass into standardized functional food ingredients is currently impeded by significant techno-economic barriers, primarily structural recalcitrance, compositional inconsistency, and the presence of toxic fermentation inhibitors. This review provides a comprehensive analysis of the synergistic application of microbial engineering and artificial intelligence (AI) to resolve these bioprocessing bottlenecks within a food-to-food closed-loop framework (as shown in the graphical abstract). We evaluate recent advances in engineering food-grade microbial chassis (e.g., Saccharomyces cerevisiae and Escherichia coli) to enhance lignocellulose degradation and stress tolerance. Concurrently, we examine the integration of AI across the entire value chain, covering deep learning-based rational enzyme design, genome-scale metabolic modeling, and intelligent process control for precision fermentation. Current evidence demonstrates that the hardware-software coupling of engineered strains and AI algorithms significantly enhances conversion efficiency and process robustness. Key findings highlight that AI-driven Design-Build-Test-Learn cycles facilitate the de novo creation of enzymes with superior kinetics and strains with adaptive stress response capabilities against toxins. Moreover, dynamic digital twin models effectively mitigate the impact of substrate variability, ensuring the batch-to-batch consistency required for food applications. We conclude that this data-driven synergistic paradigm is pivotal for establishing a resilient circular bioeconomy, enabling the reliable bioconversion of waste into high-value single-cell proteins, natural flavor additives, and sustainable packaging materials.

Artificial Intelligence

Deep learning guided programmable design of Escherichia coli core promoters from sequence architecture to strength control.

Core promoters are essential regulatory elements that control transcription initiation, but accurately predicting and designing their strength remains challenging due to complex sequence-function relationships and the limited generalizability of existing AI-based approaches. To address this, we developed a modular platform integrating rational library design, predictive modelling, and generative optimization into a closed-loop workflow for end-to-end core promoter engineering. Conserved and spacer region of core promoters exert distinct effects on transcriptional strength, with the former driving large-scale variation and the latter enabling finer gradation. Based on this insight, Mutation-Barcoding-Reverse Sequencing approach was used and constructed a synthetic promoter library comprising 112 955 variants with minimal redundancy and a 16 226-fold expression range. A Transformer-based model trained on this dataset achieved a Pearson correlation of 0.87 with experimentally measured promoter strengths. When combined with a conditional diffusion model, the system enabled de novo generation of promoter sequences with defined strengths, achieving a design-to-measurement correlation of 0.95 and maintaining high accuracy (R = 0.93) across varied sequence contexts. The designed promoters consistently preserved their intended strength gradients, demonstrating robust plug-and-play functionality. This work establishes a scalable and extensible platform (www.yudenglab.com) for deep learning-guided programmable design of Escherichia coli core promoters, enabling precise transcriptional control.

Promoter Regions, Genetic

Escherichia coli mutants deficient in the aspartate and aromatic amino acid aminotransferases.

Two new mutations are described which, together, eliminate essentially all the aminotransferase activity required for de novo biosynthesis of tyrosine, phenylalanine, and aspartic acid in a K-12 strain of Escherichia coli. One mutation, designated tyrB, lies at about 80 min on the E. coli map and inactivates the "tyrosine-repressible" tyrosine/phenylalanine aminotransferase. The second mutation, aspC, maps at about 20 min and inactivates a nonrespressible aspartate aminotransferase that also has activity on the aromatic amino acids. In ilvE- strains, which lack the branched-chain amino acid aminotransferase, the presence of either the tyrosine-repressible aminotransferase or the aspartate aminotransferase is sufficient for growth in the absence of exogenous tyrosine, phenylalanine, or aspartate; the tyrosine-repressible enzyme is also active in leucine biosynthesis. The ilvE gene product alone can reverse a phenylalanine requirement. Biochemical studies on extracts of strains carrying combinations of these aminotransferase mutations confirm the existence of two distinct enzymes with overlapping specificities for the alpha-keto acid analogues of tyrosine, phenylalanine, and aspartate. These enzymes can be distinguished by electrophoretic mobilities, by kinetic parameters using various substrates, and by a difference in tyrosine repressibility. In extracts of an ilvE- tyrB- aspC- triple mutant, no aminotransferase activity for the alpha-keto acids of tyrosine, phenylalanine, or aspartate could be detected.

Aspartate Aminotransferases

TaxTriage: an open-source metagenomic sequencing data analysis pipeline enabling putative pathogen detection.

MOTIVATION: TaxTriage is a comprehensive pathogen identification workflow designed for both short- and long-read untargeted DNA and RNA sequencing data. Combining read classification, mapping, and de novo assembly approaches, putative pathogens are identified through comparisons to curated pathogens and abundance expectations from healthy cohort data. Flexible installation options are enabled using Nextflow™ (NF), including cloud deployment via NF Tower (Seqera Platform) and local installation on a variety of systems, including standalone installations without external internet access. Final analysis summaries are compiled into an Organism Discovery Report, which lists likely pathogens and supporting data, including a custom confidence score. RESULTS: Evaluation of published in silico, clinical, and outbreak datasets identified performance comparable to alternative cloud-based processing pipelines for expected pathogen and co-infection detection with similar sensitivity and increased specificity. To support both public health and veterinary diagnostics communities, customization options have been incorporated to enable improved performance for host species of interest. AVAILABILITY AND IMPLEMENTATION: Source code for TaxTriage is freely available at https://github.com/jhuapl-bio/taxtriage. TaxTriage v2.1.1 has been archived on Zenodo at https://zenodo.org/records/17081354 to permit reproducible analysis as described in this manuscript.

Software

The biochemistry and in vitro activity of soluble factors of activated lymphocytes.

Activated lymphocytes release numerous products which are either synthesized de novo or in increased amounts; some of these products play a role in the regulation of the immune response and are designated as mediators of cellular immune reactions or lymphokines. The first lymphokine described was the macrophage migration inhibitory factor (MIF) which has been studied most extensively with regard to its chemical and biological properties. Using sensitive radiolabelling techniques and an antiserum against highly purified fractions of MIF we were able to identify several products of activated guinea pig lymphocytes with different molecular weights of 15.000, 30.000, 45.000, 60.000 which all had an isoelectric point of 5.2 and were all inhibitory to macrophage migration. It is suggested, that these molecules are oligomers of a common subunit of molecular weight 15.000. It was further shown, that molecules of the same physical-chemical and serological characteristics are produced by activated B-cells, L2C leukemia cells and growing fibroblasts, thus further substantiating earlier reports on the production of MIF by lymphoid and non-lymphoid cells. The described molecules were also shown not to contain determinants of the major histocompatibility complex and to be distinct from lymphotoxin, another lymphocyte activation product. It is concluded, that MIF is not a single molecule but rather a system of structurally related molecules. Their interaction with macrophages and possible relationships to macrophage activating factor is discussed.

Animals

Distinct patterns of de novo coding variants contribute to Tourette Syndrome etiology.

Tourette syndrome (TS) is a highly heritable childhood-onset neuropsychiatric disorder characterized by persistent motor and vocal tics. While both common and rare variants contribute to TS susceptibility, the role of rare de novo mutations (DNMs) remains incompletely characterized. Here, we report findings from the largest TS whole-exome sequencing study to date, analyzing 1,466 TS trios alongside 6,714 autism spectrum disorder (ASD) trios and 5,880 unaffected sibling controls from the Simons Simplex Collection (SSC) and SPARK cohorts. Leveraging a trio-based design across these cohorts enabled calibrated assessment of DNM burden while controlling for background mutation rates. We observed a significant exome-wide enrichment of protein-truncating DNMs in TS probands, particularly within genes intolerant to loss-of-function variation (pLI ≥ 0.9), with little contribution from damaging missense variants. Notably, TS probands did not exhibit enrichment in previously implicated ASD or developmental delay (DD) genes, but elsewhere in the genome, suggesting a distinct rare variant architecture. Using a Bayesian statistical framework that integrates both de novo and rare inherited coding variants, we identified three candidate TS risk genes with FDR ≤ 0.05: PPP5C , EXOC1 , and GXYLT1 . Literature shows that they have prior links to neurodevelopmental and psychiatric disorders. These findings reveal a rare variant burden in TS that is genetically distinguishable from ASD, underscore the importance of loss-of-function mutations in TS risk, and nominate novel candidate genes for future functional investigation.

Journal Article

AVITI sequencing of a four-generation CEPH/Utah pedigree confirms low mutation rates at homopolymer loci despite their low sequence complexity.

BACKGROUND: Short tandem repeats (STRs) and homopolymers are among the most mutable loci in the human genome. Despite their presumed mutability owing to replication slippage, homopolymer loci exhibit lower mutation rates and minimal paternal age effects compared to other STRs. This paradox questions if technical limitations, rather than biological mechanisms, explain these observations. RESULTS: We used the Element Biosciences AVITI platform to sequence the genomes of a 48-member, four-generation CEPH/Utah pedigree. As the AVITI platform reduces error rates at repetitive sequences compared to Illumina, this design enabled accurate mutation discovery at 90% of assayed homopolymers and a 1.7-fold increase in discoverable mutations compared to Illumina. We identified a median of 35 de novo homopolymer mutations per trio and a mutation rate of 5.28 &#xd7; 10-5 DNMs per locus per generation, confirming a lower rate than dinucleotides (1.94 &#xd7; 10-4). Most DNMs were single base-pair expansions or contractions. Despite comprising <1% of homopolymer loci, G/C homopolymers showed 18-fold higher mutation rates than A/T homopolymers; in contrast, the high dinucleotide mutation rate is not driven by a particular motif class. Parent-of-origin analysis revealed 78% of homopolymer mutations are paternal in origin, but no significant paternal age effect was observed. CONCLUSIONS: This study confirms that homopolymers exhibit lower mutation rates and lack strong paternal age effects compared to other STRs, likely owing to the combination of a lower propensity to form slippage-causing secondary structures and more efficient mismatch repair. Our set of high-quality mutations suggest these phenomena are biological rather than technical in nature. Finally, we demonstrate that AVITI sequencing unlocks previously intractable regions of the genome and will be a powerful tool for continued investigation of repeat mutation.

AVITI