PubMed HealthSearch

SEARCH · PubMed Health

Results for “Preprint”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Single-cell transcriptomic atlas of Alzheimer's disease middle temporal gyrus reveals region, cell type and sex specificity of gene expression with novel genetic risk for MERTK in female.

Alzheimer's disease, the most common age-related neurodegenerative disease, is closely associated with both amyloid-ß plaque and neuroinflammation. Two thirds of Alzheimer's disease patients are females and they have a higher disease risk. Moreover, women with Alzheimer's disease have more extensive brain histological changes than men along with more severe cognitive symptoms and neurodegeneration. To identify how sex difference induces structural brain changes, we performed unbiased massively parallel single nucleus RNA sequencing on Alzheimer's disease and control brains focusing on the middle temporal gyrus, a brain region strongly affected by the disease but not previously studied with these methods. We identified a subpopulation of selectively vulnerable layer 2/3 excitatory neurons that that were RORB-negative and CDH9-expressing. This vulnerability differs from that reported for other brain regions, but there was no detectable difference between male and female patterns in middle temporal gyrus samples. Disease-associated, but sex-independent, reactive astrocyte signatures were also present. In clear contrast, the microglia signatures of diseased brains differed between males and females. Combining single cell transcriptomic data with results from genome-wide association studies (GWAS), we identified MERTK genetic variation as a risk factor for Alzheimer's disease selectively in females. Taken together, our single cell dataset revealed a unique cellular-level view of sex-specific transcriptional changes in Alzheimer's disease, illuminating GWAS identification of sex-specific Alzheimer's risk genes. These data serve as a rich resource for interrogation of the molecular and cellular basis of Alzheimer's disease.

Journal Article

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids

Benchmarking large language models for genomic knowledge with GeneTuring.

Large language models (LLMs) show promise in biomedical research, but their effectiveness for genomic inquiry remains unclear. We developed GeneTuring, a benchmark consisting of 16 genomics tasks with 1,600 curated questions, and manually evaluated 48,000 answers from ten LLM configurations, including GPT-4o (via API, ChatGPT with web access, and a custom GPT setup), GPT-3.5, Claude 3.5, Gemini Advanced, GeneGPT (both slim and full), BioGPT, and BioMedLM. A custom GPT-4o configuration integrated with NCBI APIs, developed in this study as SeqSnap, achieved the best overall performance. GPT-4o with web access and GeneGPT demonstrated complementary strengths. Our findings highlight both the promise and current limitations of LLMs in genomics, and emphasize the value of combining LLMs with domain-specific tools for robust genomic intelligence. GeneTuring offers a key resource for benchmarking and improving LLMs in biomedical research.

Benchmark

Functional genomic analysis of non-canonical DNA regulatory elements of the aryl hydrocarbon receptor.

The aryl hydrocarbon receptor (AHR) is a ligand-dependent transcription factor activated by environmental toxicants like halogenated and polycyclic aromatic hydrocarbons, which then binds to DNA and regulates gene expression. AHR is implicated in numerous physiological processes, including liver and immune function, cell cycle control, oncogenesis, and metabolism. Traditionally, AHR binds a consensus DNA sequence (GCGTG), the xenobiotic response element (XRE), recruits coregulators, and modulates gene expression. Yet, recent evidence suggests AHR can also regulate gene expression via a non-consensus sequence (GGGA), termed the non-consensus XRE (NC-XRE). The prevalence and functional significance of NC-XRE motifs in the genome have remained unclear. While ChIP and reporter studies hinted at AHR-NC-XRE interactions, direct evidence for transcriptional regulation in a native context was lacking. In this study, we analyzed AHR binding to NC-XRE sequences genome-wide in mouse liver, integrating ChIP-seq and RNA-seq data to identify candidate AHR target genes containing NC-XRE motifs in their regulatory regions. We found NC-XRE motifs in 82% of AHR-bound DNA, significantly enriched compared to random regions, and present in promoters and enhancers of AHR targets. Functional genomics on the Serpine1 gene revealed that deleting NC-XRE motifs reduced TCDD-induced Serpine1 upregulation, demonstrating direct regulation. These findings provide the first direct evidence for AHR-mediated regulation via NC-XRE in a natural genomic context, advancing our understanding of AHR-bound DNA and its impact on gene expression and physiological relevance.

Journal Article

Beyond antibiotic resistance: the whiB7 transcription factor coordinates an adaptive response to alanine starvation in mycobacteria.

Pathogenic mycobacteria are a significant cause of morbidity and mortality worldwide. These bacteria are highly intrinsically drug resistant, making infections challenging to treat. The conserved whiB7 stress response is a key contributor to mycobacterial intrinsic drug resistance. Although we have a comprehensive structural and biochemical understanding of WhiB7, the complex set of signals that activate whiB7 expression remain less clear. It is believed that whiB7 expression is triggered by translational stalling in an upstream open reading frame (uORF) within the whiB7 5' leader, leading to antitermination and transcription into the downstream whiB7 ORF. To define the signals that activate whiB7, we employed a genome-wide CRISPRi epistasis screen and identified a diverse set of 150 mycobacterial genes whose inhibition results in constitutive whiB7 activation. Many of these genes encode amino acid biosynthetic enzymes, tRNAs, and tRNA synthetases, consistent with the proposed mechanism for whiB7 activation by translational stalling in the uORF. We show that the ability of the whiB7 5' regulatory region to sense amino acid starvation is determined by the coding sequence of the uORF. The uORF shows considerable sequence variation among different mycobacterial species, but it is universally and specifically enriched for alanine. Providing a potential rationalization for this enrichment, we find that while deprivation of many amino acids can activate whiB7 expression, whiB7 specifically coordinates an adaptive response to alanine starvation by engaging in a feedback loop with the alanine biosynthetic enzyme, aspC. Our results provide a holistic understanding of the biological pathways that influence whiB7 activation and reveal an extended role for the whiB7 pathway in mycobacterial physiology, beyond its canonical function in antibiotic resistance. These results have important implications for the design of combination drug treatments to avoid whiB7 activation, as well as help explain the conservation of this stress response across a wide range of pathogenic and environmental mycobacteria.

Preprint

Hybridization breaks species barriers in long-term coevolution of a cyanobacterial population.

Bacterial species often undergo rampant recombination yet maintain cohesive genomic identity. Ecological differences can generate recombination barriers between species and sustain genomic clusters in the short term. But can these forces prevent genomic mixing during long-term coevolution? Cyanobacteria in Yellowstone hot springs comprise several diverse species that have coevolved for hundreds of thousands of years, providing a rare natural experiment. By analyzing more than 300 single-cell genomes, we show that despite each species forming a distinct genomic cluster, much of the diversity within species is the result of hybridization driven by selection, which has mixed their ancestral genotypes. This widespread mixing is contrary to the prevailing view that ecological barriers can maintain cohesive bacterial species and highlights the importance of hybridization as a source of genomic diversity.

Journal Article

Quorum-sensing agr system of Staphylococcus aureus primes gene expression for protection from lethal oxidative stress.

The agr quorum-sensing system links Staphylococcus aureus metabolism to virulence, in part by increasing bacterial survival during exposure to lethal concentrations of H2O2, a crucial host defense against S. aureus. We now report that protection by agr surprisingly extends beyond post-exponential growth to the exit from stationary phase when the agr system is no longer turned on. Thus, agr can be considered a constitutive protective factor. Deletion of agr increased both respiration and fermentation but decreased ATP levels and growth, suggesting that Δagr cells assume a hyperactive metabolic state in response to reduced metabolic efficiency. As expected from increased respiratory gene expression, reactive oxygen species (ROS) accumulated more in the agr mutant than in wild-type cells, thereby explaining elevated susceptibility of Δagr strains to lethal H2O2 doses. Increased survival of wild-type agr cells during H2O2 exposure required sodA, which detoxifies superoxide. Additionally, pretreatment of S. aureus with respiration-reducing menadione protected Δagr cells from killing by H2O2. Thus, genetic deletion and pharmacologic experiments indicate that agr helps control endogenous ROS, thereby providing resilience against exogenous ROS. The long-lived "memory" of agr-mediated protection, which is uncoupled from agr activation kinetics, increased hematogenous dissemination to certain tissues during sepsis in ROS-producing, wild-type mice but not ROS-deficient (Nox2-/-) mice. These results demonstrate the importance of protection that anticipates impending ROS-mediated immune attack. The ubiquity of quorum sensing suggests that it protects many bacterial species from oxidative damage.

Staphylococcus aureus

Homologous chromosome recognition via nonspecific interactions.

In many organisms, most notably Drosophila, homologous chromosomes in somatic cells associate with each other, a phenomenon known as somatic homolog pairing. Unlike in meiosis, where homology is read out at the level of DNA sequence complementarity, somatic homolog pairing takes place without double strand breaks or strand invasion, thus requiring some other mechanism for homologs to recognize each other. Several studies have suggested a "specific button" model, in which a series of distinct regions in the genome, known as buttons, can associate with each other, presumably mediated by different proteins that bind to these different regions. Here we consider an alternative model, which we term the "button barcode" model, in which there is only one type of recognition site or adhesion button, present in many copies in the genome, each of which can associate with any of the others with equal affinity. An important component of this model is that the buttons are non-uniformly distributed, such that alignment of a chromosome with its correct homolog, compared with a non-homolog, is energetically favored; since to achieve nonhomologous alignment, chromosomes would be required to mechanically deform in order to bring their buttons into mutual register. We investigated several types of barcodes and examined their effect on pairing fidelity. We found that high fidelity homolog recognition can be achieved by arranging chromosome pairing buttons according to an actual industrial barcode used for warehouse sorting. By simulating randomly generated non-uniform button distributions, many highly effective button barcodes can be easily found, some of which achieve virtually perfect pairing fidelity. This model is consistent with existing literature on the effect of translocations of different sizes on homolog pairing. We conclude that a button barcode model can attain highly specific homolog recognition, comparable to that seen in actual cells undergoing somatic homolog pairing, without the need for specific interactions. This model may have implications for how meiotic pairing is achieved.

Preprint

Lanthanide-dependent isolation of phyllosphere methylotrophs selects for a phylogenetically conserved but metabolically diverse community.

Lanthanides have emerged as important metal cofactors for biological processes. Lanthanide-associated metabolisms are well-studied in leaf symbiont methylotrophic bacteria, which utilize reduced one-carbon compounds such as methanol for growth. Yet, the importance of lanthanides in plant-microbe interactions and on microbial physiology and colonization in plants remains poorly understood. To investigate this, 344 pink-pigmented facultative methylotrophs were isolated from soybean leaves by selecting for bacteria capable of methanol oxidation with lanthanide cofactors, but none were obligately lanthanide-dependent. Phylogenetic analyses revealed that all strains were nearly identical to each other and are part of the extorquens clade of Methylobacterium, despite variability in genome and plasmid sizes. Strain-specific identification was enabled by the higher resolution provided with rpoB compared to 16S rRNA as marker genes. Despite the low strain-level diversity, the metabolic capabilities of the collection diverged greatly. Strains encoding identical lanthanide-dependent alcohol dehydrogenases displayed significantly different growth rates and/or final ODs from each other on alcohols in the presence and absence of lanthanides. Several strains also lacked well-characterized lanthanide-associated genes thought to be important for phyllosphere colonization. Additionally, 3% of our isolates were capable of growth on sugars and 23% were capable of growth on aromatic acids, substantially expanding the range of substrates utilized by Methylobacterium extorquens in the phyllosphere. Our findings suggest that the expansion of metabolic capabilities, as well as differential usage of lanthanides and their influence on metabolism, among closely related strains point to evolution of niche partitioning strategies to promote colonization of the phyllosphere.

Journal Article

Exome-wide evidence of compound heterozygous effects across common phenotypes in the UK Biobank.

Exome-sequencing association studies have successfully linked rare protein-coding variation to risk of thousands of diseases. However, the relationship between rare deleterious compound heterozygous (CH) variation and their phenotypic impact has not been fully investigated. Here, we leverage advances in statistical phasing to accurately phase rare variants (MAF ~ 0.001%) in exome sequencing data from 175,587 UK Biobank (UKBB) participants, which we then systematically annotate to identify putatively deleterious CH coding variation. We show that 6.5% of individuals carry such damaging variants in the CH state, with 90% of variants occurring at MAF < 0.34%. Using a logistic mixed model framework, systematically accounting for relatedness, polygenic risk, nearby common variants, and rare variant burden, we investigate recessive effects in common complex diseases. We find six exome-wide significant () and 17 nominally significant () gene-trait associations. Among these, only four would have been identified without accounting for CH variation in the gene. We further incorporate age-at-diagnosis information from primary care electronic health records, to show that genetic phase influences lifetime risk of disease across 20 gene-trait combinations (FDR < 5%). Using a permutation approach, we find evidence for genetic phase contributing to disease susceptibility for a collection of gene-trait pairs, including FLG-asthma () and USH2A-visual impairment (). Taken together, we demonstrate the utility of phasing large-scale genetic sequencing cohorts for robust identification of the phenome-wide consequences of compound heterozygosity.

Preprint

Extrusion fountains are hallmarks of chromosome organization emerging upon zygotic genome activation.

The initiation of gene expression during development, known as zygotic genome activation (ZGA), is accompanied by massive changes in chromosome organization. However, the earliest events of chromosome folding and their functional roles remain unclear. Using Hi-C on zebrafish embryos, we discovered that chromosome folding begins early in development with the formation of "fountains", a novel element of chromosome organization. Emerging preferentially at enhancers, fountains exhibit an initial accumulation of cohesin, which later redistributes to CTCF sites at TAD borders. Knockouts of pioneer transcription factors driving ZGA enhancers result in the specific loss of fountains, establishing a causal link between enhancer activation and fountain formation. Polymer simulations demonstrate that fountains may arise as sites of facilitated cohesin loading, requiring two-sided but desynchronized loop extrusion, potentially caused by cohesin collisions with obstacles or internal switching. Moreover, we detected similar fountain patterns at enhancers in mouse cells. Fountains disappear upon acute cohesin depletion, as well as during mitosis, and reappear with cohesin loading in early G1. Altogether, fountains represent the first known enhancer-specific elements of chromosome organization and constitute starting points for chromosome folding during development, likely through facilitated cohesin loading.

Journal Article

Mega-Enhancer Bodies Organize Neuronal Long Genes in the Cerebellum.

Dynamic regulation of gene expression plays a key role in establishing the diverse neuronal cell types in the brain. Recent findings in genome biology suggest that three-dimensional (3D) genome organization has important, but mechanistically poorly understood functions in gene transcription. Beyond local genomic interactions between promoters and enhancers, we find that cerebellar granule neurons undergoing differentiation in vivo exhibit striking increases in long-distance genomic interactions between transcriptionally active genomic loci, which are separated by tens of megabases within a chromosome or located on different chromosomes. Among these interactions, we identify a nuclear subcompartment enriched for near-megabase long enhancers and their associated neuronal long genes encoding synaptic or signaling proteins. Neuronal long genes are differentially recruited to this enhancer-dense subcompartment to help shape the transcriptional identities of granule neuron subtypes in the cerebellum. SPRITE analyses of higher-order genomic interactions, together with IGM-based 3D genome modeling and imaging approaches, reveal that the enhancer-dense subcompartment forms prominent nuclear structures, which we term mega-enhancer bodies. These novel nuclear bodies reside in the nuclear periphery, away from other transcriptionally active structures, including nuclear speckles located in the nuclear interior. Together, our findings define additional layers of higher-order 3D genome organization closely linked to neuronal maturation and identity in the brain.

Journal Article

Oligomerization and positive feedback on membrane recruitment encode dynamically stable PAR-3 asymmetries in the C. elegans zygote.

Studies of PAR polarity have emphasized a paradigm in which mutually antagonistic PAR proteins form complementary polar domains in response to transient cues. A growing body of work suggests that the oligomeric scaffold PAR-3 can form unipolar asymmetries without mutual antagonism, but how it does so is largely unknown. Here we combine single molecule analysis and modeling to show how the interplay of two positive feedback loops promotes dynamically stable unipolar PAR-3 asymmetries in early C. elegans embryos. First, the intrinsic dynamics of PAR-3 membrane binding and oligomerization encode negative feedback on PAR-3 dissociation. Second, membrane-bound PAR-3 promotes its own recruitment through a mechanism that requires the anterior polarity proteins PAR-6 and PKC-3. Using a kinetic model tightly constrained by our experimental measurements, we show that these two feedback loops are individually required and jointly sufficient to encode dynamically stable and locally inducible unipolar PAR-3 asymmetries in the absence of posterior inhibition. Given the central role of PAR-3, and the conservation of PAR-3 membrane-binding, oligomerization, and core interactions with PAR-6/PKC-3, these results have widespread implications for PAR-mediated polarity in metazoa.

Journal Article

Active learning of enhancer and silencer regulatory grammar in photoreceptors.

Cis-regulatory elements (CREs) direct gene expression in health and disease, and models that can accurately predict their activities from DNA sequences are crucial for biomedicine. Deep learning represents one emerging strategy to model the regulatory grammar that relates CRE sequence to function. However, these models require training data on a scale that exceeds the number of CREs in the genome. We address this problem using active machine learning to iteratively train models on multiple rounds of synthetic DNA sequences assayed in live mammalian retinas. During each round of training the model actively selects sequence perturbations to assay, thereby efficiently generating informative training data. We iteratively trained a model that predicts the activities of sequences containing binding motifs for the photoreceptor transcription factor Cone-rod homeobox (CRX) using an order of magnitude less training data than current approaches. The model's internal confidence estimates of its predictions are reliable guides for designing sequences with high activity. The model correctly identified critical sequence differences between active and inactive sequences with nearly identical transcription factor binding sites, and revealed order and spacing preferences for combinations of motifs. Our results establish active learning as an effective method to train accurate deep learning models of cis-regulatory function after exhausting naturally occurring training examples in the genome.

Journal Article

Rare variation in noncoding regions with evolutionary signatures contributes to autism spectrum disorder risk.

Little is known about the role of noncoding regions in the etiology of autism spectrum disorder (ASD). We examined three classes of noncoding regions: Human Accelerated Regions (HARs), which show signatures of positive selection in humans; experimentally validated neural Vista Enhancers (VEs); and conserved regions predicted to act as neural enhancers (CNEs). Targeted and whole genome analysis of >16,600 samples and >4900 ASD probands revealed that likely recessive, rare, inherited variants in HARs, VEs, and CNEs substantially contribute to ASD risk in probands whose parents share ancestry, which enriches for recessive contributions, but modestly, if at all, in simplex family structures. We identified multiple patient variants in HARs near IL1RAPL1 and in a VE near SIM1 and showed that they change enhancer activity. Our results implicate both human-evolved and evolutionarily conserved noncoding regions in ASD risk and suggest potential mechanisms of how changes in regulatory regions can modulate social behavior.

Human Accelerated Regions

Natural variation of immune epitopes reveals intrabacterial antagonism.

Plants and animals detect biomolecules termed Microbe-Associated Molecular Patterns (MAMPs) and induce immunity. Agricultural production is severely impacted by pathogens which can be controlled by transferring immune receptors. However, most studies use a single MAMP epitope and the impact of diverse multi-copy MAMPs on immune induction is unknown. Here we characterized the epitope landscape from five proteinaceous MAMPs across 4,228 plant-associated bacterial genomes. Despite the diversity sampled, natural variation was constrained and experimentally testable. Immune perception in both Arabidopsis and tomato depended on both epitope sequence and copy number variation. For example, Elongation Factor Tu is predominantly single copy and 92% of its epitopes are immunogenic. Conversely, 99.9% of bacterial genomes contain multiple Cold Shock Proteins and 46% carry a non-immunogenic form. We uncovered a new mechanism for immune evasion, intrabacterial antagonism, where a non-immunogenic Cold Shock Protein blocks perception of immunogenic forms encoded in the same genome. These data will lay the foundation for immune receptor deployment and engineering based on natural variation.

comparative genomics

Asymmetric cortical projections to striatal direct and indirect pathways distinctly control actions.

The striatal direct and indirect pathways constitute the core for basal ganglia function in action control. Although both striatal D1- and D2-spiny projection neurons (SPNs) receive excitatory inputs from the cerebral cortex, whether or not they share inputs from the same cortical neurons, and how pathway-specific corticostriatal projections control behavior remain largely unknown. Here using a G-deleted rabies system in mice, we found that more than two-thirds of excitatory inputs to D2-SPNs also target D1-SPNs, while only one-third do so vice versa. Optogenetic stimulation of striatal D1- vs. D2-SPN-projecting cortical neurons differently regulate locomotion, reinforcement learning and sequence behavior, implying the functional dichotomy of pathway-specific corticostriatal subcircuits. These results reveal the partially segregated yet asymmetrically overlapping cortical projections on striatal D1- vs. D2-SPNs, and that the pathway-specific corticostriatal subcircuits distinctly control behavior. It has important implications in a wide range of neurological and psychiatric diseases affecting cortico-basal ganglia circuitry.

Journal Article

Plasticity of Human Microglia and Brain Perivascular Macrophages in Aging and Alzheimer's Disease.

The complex roles of myeloid cells, including microglia and perivascular macrophages, are central to the neurobiology of Alzheimer's disease (AD), yet they remain incompletely understood. Here, we profiled 832,505 human myeloid cells from the prefrontal cortex of 1,607 unique donors covering the human lifespan and varying degrees of AD neuropathology. We delineated 13 transcriptionally distinct myeloid subtypes organized into 6 subclasses and identified AD-associated adaptive changes in myeloid cells over aging and disease progression. The GPNMB subtype, linked to phagocytosis, increased significantly with AD burden and correlated with polygenic AD risk scores. By organizing AD-risk genes into a regulatory hierarchy, we identified and validated MITF as an upstream transcriptional activator of GPNMB, critical for maintaining phagocytosis. Through cell-to-cell interaction networks, we prioritized APOE-SORL1 and APOE-TREM2 ligand-receptor pairs, associated with AD progression. In both human and mouse models, TREM2 deficiency disrupted GPNMB expansion and reduced phagocytic function, suggesting that GPNMB's role in neuroprotection was TREM2-dependent. Our findings clarify myeloid subtypes implicated in aging and AD, advancing the mechanistic understanding of their role in AD and aiding therapeutic discovery.

Journal Article