PubMed HealthSearch

SEARCH · PubMed Health

Results for “Preprint”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Artificial intelligence-derived myocardial fibrosis on cardiac magnetic resonance for prognosis in cardiomyopathy: A systematic review of a sparse evidence base.

BACKGROUND: Myocardial fibrosis on cardiovascular magnetic resonance (CMR), assessed by late gadolinium enhancement (LGE) and parametric mapping, is an established predictor of adverse events in cardiomyopathy. We assessed whether artificial intelligence (AI) quantification of fibrosis adds independent prognostic value. METHODS: We searched six databases, a clinical-trials register, and a preprint server from inception to 13 June 2026. Eligible studies used AI to generate a fibrosis marker in adults with ischemic or nonischemic cardiomyopathy, with covariate-adjusted outcomes over ≥12 months. Risk of bias was assessed using PROBAST, PROBAST+AI, and QUIPS. Fewer than three comparable studies precluded meta-analysis; certainty was rated using GRADE. RESULTS: Of 448 records (381 after de-duplication), 18 full texts were reviewed and two included, one peer-reviewed and one preprint. In an ischemic-cardiomyopathy registry (Ghanbari et al.; n = 216 analytic, 26 events), AI-derived dense LGE scar predicted arrhythmic events (univariable hazard ratio [HR] 2.35, 95% CI 1.33-4.15), and AI-derived but not manual scar improved discrimination beyond guideline criteria (area under the curve 0.63 to 0.68; p = 0.02). In a nonischemic dilated-cardiomyopathy preprint (Kim et al.; n = 347, 119 events), automated extracellular volume ≥30% predicted cardiovascular death or heart-failure hospitalization (adjusted HR 2.00, 95% CI 1.32-3.03). Both were at high risk of bias, with data-derived thresholds and no external validation. CONCLUSIONS: Across only two studies, AI-derived fibrosis was independently associated with adverse cardiovascular events, but its added value over manual quantification remains unproven. Certainty was very low. The evidence base is sparse and not yet ready for clinical use.

Humans

Emerging protein sequencing technologies: proteomics without mass spectrometry?

INTRODUCTION: Liquid chromatography-tandem mass spectrometry (LC-MS/MS) has been a leading method for proteomics for 30 years. Advantages provided by LC-MS/MS are offset by significant disadvantages, including cost. Recently, several non-mass spectrometric methods have emerged, but little information is available about their capacity to analyze the complex mixtures routine for mass spectrometry. AREAS COVERED: We review recent non-mass-spectrometric methods for sequencing proteins and peptides, including those using nanopores, sequencing by degradation, reverse translation, and short-epitope mapping, with comments on bioinformatics challenges, fundamental limitations, and areas where new technologies will be more or less competitive with LC-MS/MS. In addition to conventional literature searches, instrument vendor websites, patents, webinars, and preprints were also consulted to give a more up-to-date picture. EXPERT OPINION: Many new technologies are promising. However, demonstrations that they outperform mass spectrometry in terms of peptides and proteins identified have not yet been published, and astute observers note important disadvantages, especially relating to the dynamic range of single-molecule measurements of complex mixtures. Still, even if the performance of emerging methods proves inferior to LC-MS/MS, their low cost could create a different kind of revolution: a dramatic increase in the number of biology laboratories engaging in new forms of proteomics research.

Proteomics

Lit-OTAR framework for extracting biological evidences from literature.

SUMMARY: The lit-OTAR framework, developed through a collaboration between Europe PMC and Open Targets, leverages deep learning to revolutionize drug discovery by extracting evidence from scientific literature for drug target identification and validation. This novel framework combines named entity recognition for identifying gene/protein (target), disease, organism, and chemical/drug within scientific texts, and entity normalization to map these entities to databases like Ensembl, Experimental Factor Ontology, and ChEMBL. Continuously operational, it has processed over 39 million abstracts and 4.5 million full-text articles and preprints to date, identifying more than 48.5 million unique associations that significantly help accelerate the drug discovery process and scientific research >29.9 m distinct target-disease, 11.8 m distinct target-drug, and 8.3 m distinct disease-drug relationships. AVAILABILITY AND IMPLEMENTATION: The results are accessible through Europe PMC's SciLite web app (https://europepmc.org/) and its annotations API (https://europepmc.org/annotationsapi), as well as via the Open Targets Platform (https://platform.opentargets.org/). The daily pipeline is available at https://github.com/ML4LitS/otar-maintenance, and the Open Targets ETL processes are available at https://github.com/opentargets.

Drug Discovery

The Re-Emergence of Bundibugyo Ebolavirus in Uganda and the Democratic Republic of Congo: Epidemiological Drivers, Response Strategies, and Implications for Global Health Security.

Bundibugyo ebolavirus (BDBV) is one of the least studied species within the genus Orthoebolavirus (family Filoviridae), despite its capacity to cause severe Ebola virus disease (EVD) with substantial mortality. First identified during a 2007-2008 outbreak in Bundibugyo District, western Uganda (149 reported cases, 37 deaths; case-fatality rate [CFR] approximately 25-36%), BDBV re-emerged in 2012 in Orientale Province, Democratic Republic of the Congo (DRC) (57-59 cases, 29-34 deaths; CFR 34-58%), before resurfacing in Ituri Province, DRC, in April-May 2026. By 11 August 2026, this third outbreak had grown to 4566 laboratory-confirmed cases and 2128 deaths (CFR ≈ 47%) across five DRC provinces and Uganda, becoming the largest, fastest-growing BDBV epidemic on record and the second-largest Ebola-family outbreak overall. This narrative review, not a systematic review or meta-analysis, summarizes peer-reviewed literature, preprints, and official situation reports from WHO, Africa CDC, US CDC, ECDC, and national health ministries, identified through PubMed, Scopus, Web of Science, Google Scholar, and Embase from inception to 12 August 2026, to examine BDBV historical evolution, virology and pathogenesis, drivers of re-emergence, surveillance and response, therapeutic and vaccine gaps, and global health security implications. The 2026 outbreak, unfolding amid conflict and mass displacement in eastern DRC, has been marked by an estimated basic reproduction number of 1.4-2.1 (central estimate 1.71), disproportionate infection among healthcare workers (7.2% of confirmed cases in DRC, 20% in Uganda), and the continued absence of licensed BDBV-specific vaccines or therapeutics. Findings underscore the need for sustained genomic and ecological surveillance, decentralized rapid diagnostics, broadly protective pan-filovirus vaccines, conflict-sensitive response strategies, and strengthened Uganda-DRC collaboration. Because the evidence base for the ongoing outbreak remains preliminary, findings should be interpreted cautiously and revisited as further peer-reviewed data emerge.

Bundibugyo ebolavirus

A minimal three-arm oral regimen for healthspan: mechanistic alignment with transcriptomic signals from a large parental-lifespan GWAS.

A large genome-wide association study of parental lifespan was reported in 2019. A later transcriptome-wide association study (TWAS) based on those summary statistics identified a set of transcriptional programs associated with longer genetically predicted survival, including increased brain NAD + salvage, especially NMNAT2, reduced glucose-stimulated insulin secretion, a shift toward synaptic pruning with less broad plasticity, and a glial pattern characterized by relatively greater microglial and lower astrocytic signatures, with only weak pan-tissue senescence signals. Building on those directional findings, this short communication proposes a minimal three-arm oral regimen with unequal evidentiary weight: first, the Cheung Glutamatergic Regimen, consisting of low-dose dextromethorphan potentiated by a CYP2D6 inhibitor together with piracetam and L-glutamine, as an exploratory adjunct aimed at preserving residual functional connectivity; second, daily nicotinamide mononucleotide and N-acetylcysteine with pulsed senolytics for NAD + salvage and senescence modulation; and third, GLP-1 receptor agonism for metabolic reprogramming. The NAD+/senescence arm is the primary mechanistic anchor, GLP-1 receptor agonism provides secondary metabolic support, and the glutamatergic arm is exploratory. Each arm targets a separate node within the pruning-plasticity-metabolic triad. The regimen is fully oral, uses conservative dosing, and draws on prior therapeutic or human-exposure data, although the proposed combination has no established safety profile. Although direct combination data are lacking and the foundational TWAS remains a preprint, the components show plausible but uneven mechanistic alignment with the TWAS signals and may justify carefully designed, safety-focused pilot evaluation.

GLP-1

Reconstructing the 3D genome organization of Neanderthals reveals that chromatin folding shaped phenotypic and sequence divergence.

Changes in gene regulation were a major driver of the divergence of archaic hominins (AHs)-Neanderthals and Denisovans-and modern humans (MHs). The three-dimensional (3D) folding of the genome is critical for regulating gene expression; however, its role in recent human evolution has not been explored because the degradation of ancient samples does not permit experimental determination of AH 3D genome folding. To fill this gap, we apply novel deep learning methods for inferring 3D genome organization from DNA sequence to Neanderthal, Denisovan, and diverse MH genomes. Using the resulting 3D contact maps across the genome, we identify 167 distinct regions with diverged 3D genome organization between AHs and MHs. We show that these 3D-diverged loci are enriched for genes related to the function and morphology of the eye, supra-orbital ridges, hair, lungs, immune response, and cognition. Despite these specific diverged loci, the 3D genome of AHs and MHs is more similar than expected based on sequence divergence, suggesting that the pressure to maintain 3D genome organization constrained hominin sequence evolution. We also find that 3D genome organization constrained the landscape of AH ancestry in MHs today: regions more tolerant of 3D variation are enriched for introgression in modern Eurasians. Finally, we identify loci where modern Eurasians have inherited novel 3D genome folding patterns from AH ancestors and validate folding differences in a high-frequency locus using Hi-C, revealing a putative molecular mechanism for phenotypes associated with archaic introgression. In summary, our application of deep learning to predict archaic 3D genome organization illustrates the potential of inferring molecular phenotypes from ancient DNA to reveal previously unobservable biological differences.

Journal Article

A DUAL MTOR/NAD+ ACTING GEROTHERAPY.

The geroscience hypothesis states that a therapy that prevents the underlying aging process should prevent multiple aging related diseases. The mTOR (mechanistic target of rapamycin)/insulin and NAD+ (nicotinamide adenine dinucleotide) pathways are two of the most validated aging pathways. Yet, it's largely unclear how they might talk to each other in aging. In genome-wide CRISPRa screening with a novel class of N-O-Methyl-propanamide-containing compounds we named BIOIO-1001, we identified lipid metabolism centering on SIRT3 as a point of intersection of the mTOR/insulin and NAD+ pathways. In vivo testing indicated that BIOIO-1001 reduced high fat, high sugar diet-induced metabolic derangements, inflammation, and fibrosis, each being characteristic of non-alcoholic steatohepatitis (NASH). An unbiased screen of patient datasets suggested a potential link between the anti-inflammatory and anti-fibrotic effects of BIOIO-1001 in NASH models to those in amyotrophic lateral sclerosis (ALS). Directed experiments subsequently determined that BIOIO-1001 was protective in both sporadic and familial ALS models. Both NASH and ALS have no treatments and suffer from a lack of convenient biomarkers to monitor therapeutic efficacy. A potential strength in considering BIOIO-1001 as a therapy is that the blood biomarker that it modulates, namely plasma triglycerides, can be conveniently used to screen patients for responders. More conceptually, to our knowledge BIOIO-1001 is a first therapy that fits the geroscience hypothesis by acting on multiple core aging pathways and that can alleviate multiple conditions after they have set in.

Preprint

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

Ribosomal DNA and RNA (rDNA and rRNA) sequences are usually discarded from sequencing analyses. But with hundreds of copies of rDNA genes it is unknown whether they possess sequence variations that form different types of ribosomes that affect human physiology and disease. Here, we developed an algorithm for variant-calling between paralog genes (termed RGA) and compared rDNA variations found in short- and long-read sequencing data from the 1,000 Genomes Project (1KGP) and Genome In A Bottle (GIAB). We additionally developed a novel protocol for long-read sequencing full-length rRNA (RIBO-RT) from actively translating ribosomes. Our analyses identified hundreds of rDNA variants, most of which, surprisingly, are short insertion-deletions (indels) and dozens of highly abundant rRNA variants that are incorporated into translationally active ribosomes. To visualize variant ribosomes at the single cell level, we developed an in-situ rRNA sequencing method (SWITCH-seq) which revealed that variants are co-expressed within individual cells. Strikingly, by analyzing rDNA, we found that variants assemble into distinct ribosome subtypes. We discovered that these subtypes acquire different rRNA structures by successfully employing dimethyl sulfate (DMS) probing of full length rRNA. With this atlas we investigated rRNA variation changes across human tissues and cancer types. This revealed tissue-specific rRNA subtype expression in endoderm/ectoderm-derived tissues. In cancer, low abundant rRNA variants can become highly expressed, which suggests the presence of cancer-specific ribosomes. Together, this study identifies and comprehensively characterizes the diversity of ribosomes at the level of rRNA variants which is dominated by indel variants, their chromosomal location and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Journal Article

Single-cell transcriptomic atlas of Alzheimer's disease middle temporal gyrus reveals region, cell type and sex specificity of gene expression with novel genetic risk for MERTK in female.

Alzheimer's disease, the most common age-related neurodegenerative disease, is closely associated with both amyloid-ß plaque and neuroinflammation. Two thirds of Alzheimer's disease patients are females and they have a higher disease risk. Moreover, women with Alzheimer's disease have more extensive brain histological changes than men along with more severe cognitive symptoms and neurodegeneration. To identify how sex difference induces structural brain changes, we performed unbiased massively parallel single nucleus RNA sequencing on Alzheimer's disease and control brains focusing on the middle temporal gyrus, a brain region strongly affected by the disease but not previously studied with these methods. We identified a subpopulation of selectively vulnerable layer 2/3 excitatory neurons that that were RORB-negative and CDH9-expressing. This vulnerability differs from that reported for other brain regions, but there was no detectable difference between male and female patterns in middle temporal gyrus samples. Disease-associated, but sex-independent, reactive astrocyte signatures were also present. In clear contrast, the microglia signatures of diseased brains differed between males and females. Combining single cell transcriptomic data with results from genome-wide association studies (GWAS), we identified MERTK genetic variation as a risk factor for Alzheimer's disease selectively in females. Taken together, our single cell dataset revealed a unique cellular-level view of sex-specific transcriptional changes in Alzheimer's disease, illuminating GWAS identification of sex-specific Alzheimer's risk genes. These data serve as a rich resource for interrogation of the molecular and cellular basis of Alzheimer's disease.

Journal Article

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids

Benchmarking large language models for genomic knowledge with GeneTuring.

Large language models (LLMs) show promise in biomedical research, but their effectiveness for genomic inquiry remains unclear. We developed GeneTuring, a benchmark consisting of 16 genomics tasks with 1,600 curated questions, and manually evaluated 48,000 answers from ten LLM configurations, including GPT-4o (via API, ChatGPT with web access, and a custom GPT setup), GPT-3.5, Claude 3.5, Gemini Advanced, GeneGPT (both slim and full), BioGPT, and BioMedLM. A custom GPT-4o configuration integrated with NCBI APIs, developed in this study as SeqSnap, achieved the best overall performance. GPT-4o with web access and GeneGPT demonstrated complementary strengths. Our findings highlight both the promise and current limitations of LLMs in genomics, and emphasize the value of combining LLMs with domain-specific tools for robust genomic intelligence. GeneTuring offers a key resource for benchmarking and improving LLMs in biomedical research.

Benchmark

Functional genomic analysis of non-canonical DNA regulatory elements of the aryl hydrocarbon receptor.

The aryl hydrocarbon receptor (AHR) is a ligand-dependent transcription factor activated by environmental toxicants like halogenated and polycyclic aromatic hydrocarbons, which then binds to DNA and regulates gene expression. AHR is implicated in numerous physiological processes, including liver and immune function, cell cycle control, oncogenesis, and metabolism. Traditionally, AHR binds a consensus DNA sequence (GCGTG), the xenobiotic response element (XRE), recruits coregulators, and modulates gene expression. Yet, recent evidence suggests AHR can also regulate gene expression via a non-consensus sequence (GGGA), termed the non-consensus XRE (NC-XRE). The prevalence and functional significance of NC-XRE motifs in the genome have remained unclear. While ChIP and reporter studies hinted at AHR-NC-XRE interactions, direct evidence for transcriptional regulation in a native context was lacking. In this study, we analyzed AHR binding to NC-XRE sequences genome-wide in mouse liver, integrating ChIP-seq and RNA-seq data to identify candidate AHR target genes containing NC-XRE motifs in their regulatory regions. We found NC-XRE motifs in 82% of AHR-bound DNA, significantly enriched compared to random regions, and present in promoters and enhancers of AHR targets. Functional genomics on the Serpine1 gene revealed that deleting NC-XRE motifs reduced TCDD-induced Serpine1 upregulation, demonstrating direct regulation. These findings provide the first direct evidence for AHR-mediated regulation via NC-XRE in a natural genomic context, advancing our understanding of AHR-bound DNA and its impact on gene expression and physiological relevance.

Journal Article

Beyond antibiotic resistance: the whiB7 transcription factor coordinates an adaptive response to alanine starvation in mycobacteria.

Pathogenic mycobacteria are a significant cause of morbidity and mortality worldwide. These bacteria are highly intrinsically drug resistant, making infections challenging to treat. The conserved whiB7 stress response is a key contributor to mycobacterial intrinsic drug resistance. Although we have a comprehensive structural and biochemical understanding of WhiB7, the complex set of signals that activate whiB7 expression remain less clear. It is believed that whiB7 expression is triggered by translational stalling in an upstream open reading frame (uORF) within the whiB7 5' leader, leading to antitermination and transcription into the downstream whiB7 ORF. To define the signals that activate whiB7, we employed a genome-wide CRISPRi epistasis screen and identified a diverse set of 150 mycobacterial genes whose inhibition results in constitutive whiB7 activation. Many of these genes encode amino acid biosynthetic enzymes, tRNAs, and tRNA synthetases, consistent with the proposed mechanism for whiB7 activation by translational stalling in the uORF. We show that the ability of the whiB7 5' regulatory region to sense amino acid starvation is determined by the coding sequence of the uORF. The uORF shows considerable sequence variation among different mycobacterial species, but it is universally and specifically enriched for alanine. Providing a potential rationalization for this enrichment, we find that while deprivation of many amino acids can activate whiB7 expression, whiB7 specifically coordinates an adaptive response to alanine starvation by engaging in a feedback loop with the alanine biosynthetic enzyme, aspC. Our results provide a holistic understanding of the biological pathways that influence whiB7 activation and reveal an extended role for the whiB7 pathway in mycobacterial physiology, beyond its canonical function in antibiotic resistance. These results have important implications for the design of combination drug treatments to avoid whiB7 activation, as well as help explain the conservation of this stress response across a wide range of pathogenic and environmental mycobacteria.

Preprint

Hybridization breaks species barriers in long-term coevolution of a cyanobacterial population.

Bacterial species often undergo rampant recombination yet maintain cohesive genomic identity. Ecological differences can generate recombination barriers between species and sustain genomic clusters in the short term. But can these forces prevent genomic mixing during long-term coevolution? Cyanobacteria in Yellowstone hot springs comprise several diverse species that have coevolved for hundreds of thousands of years, providing a rare natural experiment. By analyzing more than 300 single-cell genomes, we show that despite each species forming a distinct genomic cluster, much of the diversity within species is the result of hybridization driven by selection, which has mixed their ancestral genotypes. This widespread mixing is contrary to the prevailing view that ecological barriers can maintain cohesive bacterial species and highlights the importance of hybridization as a source of genomic diversity.

Journal Article

Quorum-sensing agr system of Staphylococcus aureus primes gene expression for protection from lethal oxidative stress.

The agr quorum-sensing system links Staphylococcus aureus metabolism to virulence, in part by increasing bacterial survival during exposure to lethal concentrations of H2O2, a crucial host defense against S. aureus. We now report that protection by agr surprisingly extends beyond post-exponential growth to the exit from stationary phase when the agr system is no longer turned on. Thus, agr can be considered a constitutive protective factor. Deletion of agr increased both respiration and fermentation but decreased ATP levels and growth, suggesting that Δagr cells assume a hyperactive metabolic state in response to reduced metabolic efficiency. As expected from increased respiratory gene expression, reactive oxygen species (ROS) accumulated more in the agr mutant than in wild-type cells, thereby explaining elevated susceptibility of Δagr strains to lethal H2O2 doses. Increased survival of wild-type agr cells during H2O2 exposure required sodA, which detoxifies superoxide. Additionally, pretreatment of S. aureus with respiration-reducing menadione protected Δagr cells from killing by H2O2. Thus, genetic deletion and pharmacologic experiments indicate that agr helps control endogenous ROS, thereby providing resilience against exogenous ROS. The long-lived "memory" of agr-mediated protection, which is uncoupled from agr activation kinetics, increased hematogenous dissemination to certain tissues during sepsis in ROS-producing, wild-type mice but not ROS-deficient (Nox2-/-) mice. These results demonstrate the importance of protection that anticipates impending ROS-mediated immune attack. The ubiquity of quorum sensing suggests that it protects many bacterial species from oxidative damage.

Staphylococcus aureus

Homologous chromosome recognition via nonspecific interactions.

In many organisms, most notably Drosophila, homologous chromosomes in somatic cells associate with each other, a phenomenon known as somatic homolog pairing. Unlike in meiosis, where homology is read out at the level of DNA sequence complementarity, somatic homolog pairing takes place without double strand breaks or strand invasion, thus requiring some other mechanism for homologs to recognize each other. Several studies have suggested a "specific button" model, in which a series of distinct regions in the genome, known as buttons, can associate with each other, presumably mediated by different proteins that bind to these different regions. Here we consider an alternative model, which we term the "button barcode" model, in which there is only one type of recognition site or adhesion button, present in many copies in the genome, each of which can associate with any of the others with equal affinity. An important component of this model is that the buttons are non-uniformly distributed, such that alignment of a chromosome with its correct homolog, compared with a non-homolog, is energetically favored; since to achieve nonhomologous alignment, chromosomes would be required to mechanically deform in order to bring their buttons into mutual register. We investigated several types of barcodes and examined their effect on pairing fidelity. We found that high fidelity homolog recognition can be achieved by arranging chromosome pairing buttons according to an actual industrial barcode used for warehouse sorting. By simulating randomly generated non-uniform button distributions, many highly effective button barcodes can be easily found, some of which achieve virtually perfect pairing fidelity. This model is consistent with existing literature on the effect of translocations of different sizes on homolog pairing. We conclude that a button barcode model can attain highly specific homolog recognition, comparable to that seen in actual cells undergoing somatic homolog pairing, without the need for specific interactions. This model may have implications for how meiotic pairing is achieved.

Preprint

Lanthanide-dependent isolation of phyllosphere methylotrophs selects for a phylogenetically conserved but metabolically diverse community.

Lanthanides have emerged as important metal cofactors for biological processes. Lanthanide-associated metabolisms are well-studied in leaf symbiont methylotrophic bacteria, which utilize reduced one-carbon compounds such as methanol for growth. Yet, the importance of lanthanides in plant-microbe interactions and on microbial physiology and colonization in plants remains poorly understood. To investigate this, 344 pink-pigmented facultative methylotrophs were isolated from soybean leaves by selecting for bacteria capable of methanol oxidation with lanthanide cofactors, but none were obligately lanthanide-dependent. Phylogenetic analyses revealed that all strains were nearly identical to each other and are part of the extorquens clade of Methylobacterium, despite variability in genome and plasmid sizes. Strain-specific identification was enabled by the higher resolution provided with rpoB compared to 16S rRNA as marker genes. Despite the low strain-level diversity, the metabolic capabilities of the collection diverged greatly. Strains encoding identical lanthanide-dependent alcohol dehydrogenases displayed significantly different growth rates and/or final ODs from each other on alcohols in the presence and absence of lanthanides. Several strains also lacked well-characterized lanthanide-associated genes thought to be important for phyllosphere colonization. Additionally, 3% of our isolates were capable of growth on sugars and 23% were capable of growth on aromatic acids, substantially expanding the range of substrates utilized by Methylobacterium extorquens in the phyllosphere. Our findings suggest that the expansion of metabolic capabilities, as well as differential usage of lanthanides and their influence on metabolism, among closely related strains point to evolution of niche partitioning strategies to promote colonization of the phyllosphere.

Journal Article

Exome-wide evidence of compound heterozygous effects across common phenotypes in the UK Biobank.

Exome-sequencing association studies have successfully linked rare protein-coding variation to risk of thousands of diseases. However, the relationship between rare deleterious compound heterozygous (CH) variation and their phenotypic impact has not been fully investigated. Here, we leverage advances in statistical phasing to accurately phase rare variants (MAF ~ 0.001%) in exome sequencing data from 175,587 UK Biobank (UKBB) participants, which we then systematically annotate to identify putatively deleterious CH coding variation. We show that 6.5% of individuals carry such damaging variants in the CH state, with 90% of variants occurring at MAF < 0.34%. Using a logistic mixed model framework, systematically accounting for relatedness, polygenic risk, nearby common variants, and rare variant burden, we investigate recessive effects in common complex diseases. We find six exome-wide significant () and 17 nominally significant () gene-trait associations. Among these, only four would have been identified without accounting for CH variation in the gene. We further incorporate age-at-diagnosis information from primary care electronic health records, to show that genetic phase influences lifetime risk of disease across 20 gene-trait combinations (FDR < 5%). Using a permutation approach, we find evidence for genetic phase contributing to disease susceptibility for a collection of gene-trait pairs, including FLG-asthma () and USH2A-visual impairment (). Taken together, we demonstrate the utility of phasing large-scale genetic sequencing cohorts for robust identification of the phenome-wide consequences of compound heterozygosity.

Preprint