PubMed HealthSearch

SEARCH · PubMed Health

Results for “big genomes”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Interpreting cancer genetics through a two-step "evolutionary cascade hypothesis": bridging neutral and selective perspectives.

BACKGROUND: DNA mutations are the fundamental engines of cancer, driving its initiation and progression. The forces that fuel malignancy are also the architects of evolution, shaping life through genetic variations. Mutations, in fact, can emerge naturally from endogenous processes, such as oxidative DNA damage or errors in replication, as well as induced by external factors, including cosmic radiation and chemical carcinogens. MAIN BODY: A key question in cancer research is whether tumor evolution is primarily governed by selective bottlenecks, neutral evolution, or dynamic genetic plasticity. In this work, we examine cancer as a disease driven by evolutionary processes rooted in fundamental biological requirements, including sustained proliferation and nutrient utilization. We hypothesize that the accumulation of mutations activates an evolutionary switch, enabling tumor cells to acquire an enhanced capacity for survival, adaptation, and growth at rates far exceeding typical evolutionary timescales. We propose the "evolutionary cascade hypothesis," a unifying framework that integrates these models into a coherent sequence. At its core lies the failure of DNA repair mechanisms, representing a critical transition in cancer progression. This shift marks the transition from an initial non-Darwinian, neutral phase to a Darwinian, more deterministic phase. CONCLUSIONS: As predictive models of tumor evolution advance through genomic big data and artificial intelligence-driven analysis, the future of cancer treatment may extend beyond targeting individual mutations to disrupting the underlying evolutionary mechanisms that sustain malignancy. This paradigm shift could redefine therapeutic strategies and ultimately improve patient outcomes.

Humans

A century of research on the Planctomycetota bacterial phylum, previously known as Planctomycetes.

One hundred years after planctomycetes were discovered and 50 years since the first isolate was successfully cultured, this bacterial phylum remains enigmatic in many ways. In the last few decades, a significant effort to characterize new isolates has resulted in >150 described species, allowing a more comprehensive analysis of their features. However, metagenomic studies reveal that a diverse group of planctomycetes has yet to be cultured and characterized, and that many biological surprises are yet to be revealed. This is the case for the recently discovered phagotrophic Candidatus Uabimicrobium, which challenges our understanding of the distinction between prokaryotes and eukaryotes. The unique biology of planctomycete cells, such as their ability to divide without the FtsZ protein, their complex structure and characteristic morphology, their relatively large genomes containing many genes with unknown function, and their variable metabolic capabilities, imposes significant barriers for researchers. Although ubiquitous, the precise ecological roles of planctomycetes in various environments are still not fully understood. However, their distinctive metabolism opens the door to a large number of potential biotechnological applications, which are beginning to be unveiled. In this article, we first review the historical milestones in planctomycetes research and describe the pioneers of the field. We then describe the controversies and their resolutions, we highlight the past discoveries and current interrogations related to planctomycetes, and discuss the ongoing challenges that hinder a comprehensive understanding of their biology. We end up with directions for exploring the biology and ecological roles of these fascinating organisms.

Bacteria

Big data analytics for CLEC5A dynamics based on single cell genomics and proteomics reveal its diverse functions in human diseases.

BACKGROUND: CLEC5A (C-type lectin domain family 5 member A) is an innate immune receptor implicated in inflammatory signaling, contributing to hyperinflammatory responses in infections and sterile inflammation. However, CLEC5A dynamics in human diseases remain to be identified. Here, we systematically characterized CLEC5A dynamics in humans across cells, tissues, and disease states, and to explore the functional significance of CLEC5A in macrophage activation based on single-cell genomics. METHODS: With multi-omics (scRNA-seq, proteomics and big data analytics), we analyzed extensive human transcriptomic datasets (>42,000 samples) to profile CLEC5A expression by cell type, tissue, and disease. Single-nucleus RNA-seq (snRNA-seq) from pediatric congenital heart disease and a virtual CLEC5A gene knockout were also performed to characterize CLEC5A dynamics in humans. RESULTS: CLEC5A is highly enriched in innate immune cells, particularly in macrophages and neutrophils. Baseline CLEC5A in most tissues is low, but it is markedly upregulated in inflammatory and infectious diseases. CLEC5A expression has sex-specific differences in certain organs. Single-cell analysis showed that CLEC5A can be considered novel marker of proinflammatory macrophages with elevated cytokine production, antigen presentation, and impaired phagocytosis. Virtual CLEC5A knockout analysis identified coordinated perturbation of immune-regulatory pathways and overlapping genes linking CLEC5A to macrophage activation networks. CONCLUSION: CLEC5A is predominantly expressed in myeloid cells and acts as a key amplifier of inflammation in human diseases. Our findings highlight CLEC5A as a potential biomarker and therapeutic target in myeloid-driven hyperinflammatory conditions, warranting further experimental and translational validation.

Humans

Evolution and applications of genome-scale metabolic models in yeast systems biology studies.

Genome-scale metabolic models (GEMs) can be used to simulate the metabolic network of an organism in a systematic and holistic way. Different yeast species, including Saccharomyces cerevisiae, have emerged as powerful cell factories for bioproduction. Recently, with the dedicated efforts from the scientific community, significant progress has been made in the development of yeast GEMs. Numerous versions of yeast GEMs and the derived multiscale models have been released, facilitating integrative omics analysis and rational strain design for different types of yeast cell factories. These advancements reflected the evolution and maturation of yeast GEMs together with a model ecosystem around them. This review will summarize the development and expansion of yeast GEMs and discuss their applications in yeast systems biology studies. It is anticipated that yeast GEMs will continue to play an increasingly important role in pioneering yeast physiological and metabolic studies in coming years.

Systems Biology

Personality Genomics.

Recent research advances have precipitated the era of personality genomics: the study of how variation in human DNA sequence predicts individual differences in characteristic patterns of thinking, feeling, and behaving. Here, we introduce personality and genomics, and we review key findings from recent genome-wide association studies of personality traits. These findings support five key observations: (a) sizable genetic effects on personality arise from a vast number of genetic variants with individually miniscule effects; (b) genetic variants associated with personality have widespread associations with other attributes, including social, economic, and medical outcomes; (c) genetic effects on personality generalize across groupings of people; (d) genetic effects on personality are minimally confounded by familial environmental effects; and (e) many recent genomic findings were anticipated by classic twin genetic research. For personality psychologists, embracing genomics provides unique and powerful inferential tools. For genomics researchers, incorporating unifying personality frameworks enables an integrative understanding of core behavioral dimensions.

Humans

Dissemination of antimicrobial resistance in Klebsiella spp. from urban aquatic environments: a multi-country genomic perspective.

INTRODUCTION: Antibiotic resistance, particularly carbapenem-resistant Klebsiella pneumoniae (CRKP), poses significant clinical and environmental threats, especially in urban aquatic ecosystems and hospital wastewaters. OBJECTIVES: This study aims to analyze the epidemiological and genomic features of CRKP isolates in urban aquatic environments and evaluate their public health and environmental impacts. METHODS AND RESULTS: Water samples were collected from 113 rivers and 3 hospitals in China, Sri Lanka, and Nepal to isolate carbapenem-resistant Klebsiella spp. isolates. Antimicrobial susceptibility testing, whole-genome sequencing, and bioinformatics analyses were performed to characterize resistance phenotypes, antibiotic resistance genes (ARGs), and evolutionary trends. Big data analysis further elucidated the genomic characteristics of CRKP in global water sources, and Galleria mellonella larvae were used to assess virulence. Statistical analysis validated the findings. A total of 192 carbapenem-resistant Klebsiella spp. isolates were identified from urban aquatic ecosystems in China (n = 60) and Nepal (n = 132), with CRKP (n = 161) being the predominant species. All CRKP isolates exhibited a multidrug-resistant phenotype, yet significant differences in resistance profiles and associated ARGs were observed between isolates from the two countries. Nine carbapenem resistance genes (CRGs) were detected, with blaNDM-1 being the most prevalent (57.8 %). Correlation analysis revealed a strong association between these CRGs and multiple Inc-type plasmids. Global genomic analysis of CRKP from water sources across eight countries identified ten distinct CRGs across 45 serotypes, with KL64 being the most predominant. Notably, carbapenem-resistant hypervirulent Klebsiella pneumoniae was detected in water samples from Nepal. CONCLUSION: Our findings highlight significant regional disparities in CRKP prevalence and ARG dissemination across urban aquatic environments, with Nepal showing the highest prevalence, particularly in untreated rivers. China exhibited lower prevalence but distinct resistance gene profiles, while no CRKP was detected in Sri Lanka, underscoring the impact of environmental management and healthcare infrastructure on ARG spread.

Humans

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans

Big data in multiple sclerosis.

PURPOSE OF REVIEW: This review summarizes recent key advancements in multiple sclerosis (MS) achieved through the utilization of big data from diverse sources and advanced analytical techniques. RECENT FINDINGS: Real-world evidence (RWE) derived from MS big data has significantly enhanced treatment strategies, redefined the concept of disease progression, refined prognostic models, and facilitated personalized medicine. RWE has highlighted the long-term benefits of early intensive treatment compared to escalation strategies, the unfavorable risk profile associated with treatment de-escalation and the importance of managing treatments during pregnancy. Additionally, it has revealed similarities and differences in the effectiveness and safety of specific high-efficacy therapies, as well as key predictors for switching treatments. RWE has also emphasized the central role of progression independent of relapse activity as a significant driver of disability and predictor of unfavorable long-term outcomes in both adult and pediatric onset MS. A data-driven approach utilizing artificial intelligence and big data has established a comprehensive framework for understanding the disease's evolution. Multimodal big data frameworks - encompassing clinical data, MRI, genomics, biomarkers, and app-based metrics - have demonstrated their ability to enhance diagnostic performance and risk stratification in MS. SUMMARY: Big data approaches are transforming MS research and clinical practice by providing stronger RWE to guide therapeutic decision-making, refining models of disease progression, and developing more precise prognostic tools.

Humans

Microbial partnerships and molecular mechanisms in plant stress physiology for climate-resilient and sustainable farming.

Plant-microbial partnerships and their underlying molecular mechanisms are indispensable, natural drivers of improved nutrient acquisition and stress tolerance in the face of climate-driven environmental challenges. Modern multi-omics tools, when coupled with artificial intelligence and synthetic biology, enable the precise design of targeted bioinoculants and synthetic microbial consortia. Translating these advanced microbiome-based strategies into scalable, field-level agricultural applications provides a sustainable path toward securing global food production while maintaining soil health. Global climate change imposes multifaceted abiotic and biotic stresses on crops, disrupting physiological and molecular processes and threatening agricultural productivity. Plant-associated microbes represent an underexplored yet powerful ally in enhancing crop resilience. This review presents current knowledge of plant-microbe interactions and the molecular mechanisms governing plant stress physiology, with an emphasis on climate-resilient and sustainable farming. Hence, ever-changing environmental cues pose a significant burden on agricultural productivity, and plant-associated microbial communities modulate a cascade of physiological and molecular responses, including production of phytohormones, signaling, regulation of reactive oxygen species homeostasis, and activation of plant immune responses to help plants withstand stress and enhance productivity. Moreover, root exudates, phytohormones, and quorum sensing mediate the central communication networks, facilitating plant-microbe cross talk. Additionally, the advances in OMICs approaches aid in disentangling the molecular underpinnings of these interactions by providing mechanistic insights and potential candidate gene targets for crop improvement and stress resilience. In the post-genomic era, integrating artificial intelligence and big data analysis to optimize microbiome-based strategies for sustainable agriculture is a new frontier for disentangling plant-microbe symbiosis to improve soil health, enhance crop yields, and improve stress tolerance. Thus, by integrating the ecological, physiological, and molecular perspectives, this review highlights the transformative potential of harnessing plant-microbe symbiosis for climate-resilient and sustainable agriculture.

Stress, Physiological

Robust inference and correlates from genetic associations with personality.

Personality traits describe stable differences in how people think, feel and behave, and how they interact with and experience their social and physical environments1,2. Many questions remain unanswered about associations between DNA and personality traits, such as their robustness, their generalizability and the biological and social pathways through which they act. Here we meta-analyse data across 46 cohorts comprising 611,037 to 1.14 million participants with European-like and African-like genomes for genome-wide association studies (GWAS) of the Big Five personality traits (extraversion, agreeableness, conscientiousness, neuroticism and openness to experience), and data from up to 50,725 participants for within-family GWAS. We identify 1,260 lead genetic variants associated with personality, including 824 novel variants3. Common genetic variants explain a moderate 4.8-9.3% of the variance in measures of each trait, and 9.3-13.3% among instruments with typical measurement reliability. Genetic associations with personality are highly consistent but not identical across geography, reporter (self versus close other), age group and measurement instrument, and we find minimal spousal assortment for personality in recent history. In contrast to many other social and behavioural traits4,5, within-family GWAS and polygenic index analyses indicate that genetic associations with personality are minimally confounded by the shared family environment. Polygenic prediction, genetic correlation and Mendelian randomization analyses indicate that personality traits have widespread, potentially causal associations with consequential behaviours and life outcomes. Overall, we find that the genetic architecture of personality is robustly generalizable, minimally confounded and widely relevant to human experience.

Journal Article

Through the lens of bioenergy crops: advances, bottlenecks, and promises of plant engineering.

Advances in engineering of bioenergy crops were driven over the past years by adapting technological breakthroughs and accelerating conventional applications but also exposed intriguing challenges. New tools revealed rich interconnectivity in the exponentially growing and dynamic 'big' omics data' of metabolomes, transcriptomes, and genomes at previously inaccessible magnitude (global, cross-species, meta-) and resolution (single cell). Insights enabled fresh hypotheses and stimulated disciplines such as functional genomics with discovery of broad regulatory networks and their determinants, that is, DNA parts, including promoters, regulatory elements, and transcription factors. Their rational design, assembly into increasingly complex blueprints, and installation into diverse chassis is an existing frontier that may benefit from emerging technologies to address bottlenecks. Interweaving nature-inspired to fully synthetic parts has already allowed building of fine-tuned regulatory circuits, or new-to-nature metabolic routes insulated from the biological context of the chassis species. Similarly, developments and the evolving need for unifying principles in plant transformation and species-agnostic technologies highlight future opportunities for engineering the next generation of bioenergy plants.

Crops, Agricultural

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework to address this challenge, which reduces runtime and memory requirements by orders of magnitude. In PSU, small subsamples of sites from a concatenated alignment are analyzed, which are expanded by upsampling before inference, and the resulting inferences are aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-alignment analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis of concatenated alignments. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical. We suggest that PSU is a general approach for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

Phylogeny

Advances in tumor subclone formation and mechanisms of growth and invasion.

Tumor subclones refer to distinct cell populations within the same tumor that possess different genetic characteristics. They play a crucial role in understanding tumor heterogeneity, evolution, and therapeutic resistance. The formation of tumor subclones is driven by several key mechanisms, including the inherent genetic instability of tumor cells, which facilitates the accumulation of novel mutations; selective pressures from the tumor microenvironment and therapeutic interventions, which promote the expansion of certain subclones; and epigenetic modifications, such as DNA methylation and histone modifications, which alter gene expression patterns. Major methodologies for studying tumor subclones include single-cell sequencing, liquid biopsy, and spatial transcriptomics, which provide insights into clonal architecture and dynamic evolution. Beyond their direct involvement in tumor growth and invasion, subclones significantly contribute to tumor heterogeneity, immune evasion, and treatment resistance. Thus, an in-depth investigation of tumor subclones not only aids in guiding personalized precision therapy, overcoming drug resistance, and identifying novel therapeutic targets, but also enhances our ability to predict recurrence and metastasis risks while elucidating the mechanisms underlying tumor heterogeneity. The integration of artificial intelligence, big data analytics, and multi-omics technologies is expected to further advance research in tumor subclones, paving the way for novel strategies in cancer diagnosis and treatment. This review aims to provide a comprehensive overview of tumor subclone formation mechanisms, evolutionary models, analytical methods, and clinical implications, offering insights into precision oncology and future translational research.

Humans

The Computational Revolution in Natural Product Research: A Data-Driven Roadmap for Next-Generation Drug Development.

Natural products (NPs) have historically provided the foundational scaffolds for drug development, yet traditional bioprospecting faces critical limitations: high rediscovery rates, laborious isolation workflows, and substantial attrition during clinical translation. The emergence of big data technologies is fundamentally transforming this landscape, enabling a shift from serendipity-based discovery toward systematic, data-driven approaches. This review examines how the integration of artificial intelligence (AI), machine learning (ML), and multi-omics datasets is accelerating natural product research across three key domains: (1) genome mining for biosynthetic gene cluster identification using platforms such as antiSMASH, (2) cheminformatics-driven prediction of structure-activity relationships and ADMET properties, and (3) metabolomics-guided dereplication to prioritize novel bioactive scaffolds. We evaluate the convergence of genomics, metabolomics, and computational chemistry in enabling in silico lead optimization and the discovery of cryptic metabolites from previously inaccessible microbial taxa. While challenges in data standardization and scalability persist, the synergy between big data and NP research is accelerating clinical translation. Despite persistent challenges in data standardization, scalability, and equitable benefit-sharing, the convergence of big data and NP research is poised to redefine drug development. These advances position computational NP research as a cornerstone of next-generation drug development.

big data analytics

Big data and psychiatry: advances, constraints and future directions.

Early work in psychiatry research, often involving single sites, small samples, and limited variables, has shifted to contemporary research involving multiple sites, large samples, and many variables. Such research raises important questions, including concerns about data quality and methodological rigor, uncertainty about its key lessons, issues regarding clinical relevance, and questions about how to optimize future advances. Here we consider these questions and concerns against the context of big data work on community and register-based surveys, cohort and biobank studies, electronic health records, digital phenotyping, brain imaging, genomics and other -omics, and randomized controlled trials. The development of large datasets allowing well-powered analyses is a major milestone, but sample size alone does not guarantee more precise estimates, and ongoing attention to the quality and rigor of big data collation and analysis is needed. Big data research has fostered trans-disciplinarity and given insights into mechanisms underlying psychiatric disorders, but also emphasizes the intricacy, heterogeneity and variability of such mechanisms, and the importance of triangulating between large-scale and small-scale research. The complexity of psychiatric phenotypes and psychobiological mechanisms contributes to the difficulty in bridging from big data to clinical application; big data research reinforces the importance of holding our diagnoses of psychiatric disorders lightly and providing explanations of these conditions humbly; and future work needs to be more attentive to clinical issues. There is enormous scope for further building databases relevant to psychiatry, but advances in conceptual models and asking the right questions are equally valuable. The full impact of big data, including artificial intelligence analyses, remains to be seen, but overenthusiastic support should be tempered by a better understanding of its strengths and limitations. At its best, such work will contribute in an iterative and integrative way to advancing our knowledge of psychiatric disorders and mental health.

Big data