PubMed HealthSearch

SEARCH · PubMed Health

Results for “Health data integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Co-location of services: an umbrella review to consider how primary care estates could be better used to support disadvantaged groups.

AIM: To examine how co-located community and health services in primary care could support disadvantaged groups. BACKGROUND: Co-locating services is thought to improve access, collaboration, and patient outcomes. There are thousands of primary care premises across the UK. At a time of stagnating or widening health inequalities, they present an ideal opportunity to support communities, especially in disadvantaged areas. METHOD: We conducted a systematic umbrella review. Articles were retrieved from Ovid MEDLINE and Ovid Embase with supplementary snowball and grey literature searches. Reviews of co-located services supporting disadvantaged groups in primary care between 2010 and February 2024 were included. Quality and risk of bias were assessed using the Joanna Briggs Institute checklist. Two reviewers assessed eligibility, extracted data and assessed quality. Outcomes relating to health, welfare, healthcare utilization, and activity and processes were assessed. Data were narratively synthesized using a convergent integrated approach. FINDINGS: 2626 studies were screened, supplemented by snowball and grey literatures searches. Thirteen reviews were included for synthesis. One review included meta-analysis. Three models of care were identified; legal advice, welfare advice, and complementary health care. Data were synthesized according to themes: access and engagement, quality of care, efficiency, improved health, and improved social factors. We found co-located services can improve access to care, engagement in treatment, and quality of care for disadvantaged groups. Improvements to social determinants of health and mental health and well-being outcomes were reported. Findings were inconsistent when considering the impact of co-location on efficiency. We conclude that co-located services in primary care have the potential to improve identification of people most in need and improve their access to high quality health care and social support. Policy makers and practitioners should maximize the use of primary care estates to support disadvantaged groups and communities.

Humans

Advancing translational exposomics: bridging genome, exposome and personalized medicine.

Understanding the interplay between genetic predisposition and environmental and lifestyle exposures is essential for advancing precision medicine and public health. The exposome, defined as the sum of all environmental exposures an individual encounters throughout their lifetime, complements genomic data by elucidating how external and internal exposure factors influence health outcomes. This treatise highlights the emerging discipline of translational exposomics that integrates exposomics and genomics, offering a comprehensive approach to decipher the complex relationships between environmental and lifestyle exposures, genetic variability, and disease phenotypes. We highlight cutting-edge methodologies, including multi-omics technologies, exposome-wide association studies (EWAS), physiology-based biokinetic modeling, and advanced bioinformatics approaches. These tools enable precise characterization of both the external and the internal exposome, facilitating the identification of biomarkers, exposure-response relationships, and disease prediction and mechanisms. We also consider the importance of addressing socio-economic, demographic, and gender disparities in environmental health research. We emphasize how exposome data can contextualize genomic variation and enhance causal inference, especially in studies of vulnerable populations and complex diseases. By showcasing concrete examples and proposing integrative platforms for translational exposomics, this work underscores the critical need to bridge genomics and exposomics to enable precision prevention, risk stratification, and public health decision-making. This integrative approach offers a new paradigm for understanding health and disease beyond genetics alone.

Humans

The SARS-CoV-2 Integrated Genomic Epidemiology Database (IGED): Linking viral genomes with patient-level metadata to advance statewide genomic surveillance in California.

In July 2021, the California Code of Regulations Title 17 required all laboratories performing SARS‑CoV‑2 whole genome sequencing (WGS) to report their sequencing results to the California Department of Public Health (CDPH). These viral genomic data and patient metadata were compiled into the Integrated Genomic Epidemiology Database (IGED). Linking anonymized viral sequences with patient‑level information enabled monitoring of infectiousness, pathogenicity, transmission dynamics, evolution, and vaccine evasion among emerging SARS‑CoV‑2 lineages. Laboratories performing SARS-CoV-2 WGS transmitted sequencing results to CDPH through Electronic Laboratory Reporting (ELR) and non-ELR pathways. CDPH applied uniform reporting requirements but allowed flexibility in specific data formats to accommodate diverse data systems. To preserve data quality and interoperability across heterogeneous sources, CDPH implemented standardization, validation, and deduplication protocols. Snowflake, a cloud‑based data storage and analytics platform, and Posit Connect, a cloud deployment and automation platform, supported the management, processing, and integration of data within the IGED. The IGED established links between SARS‑CoV‑2 WGS data and epidemiologic metadata for 801,418 sequences, representing 81.7% of all sequences reported in California. Lineages reported to the IGED showed strong concordance with lineage proportions in GISAID. Sequences reported to the IGED had average turnaround times longer than one month, and the majority of sequencing was performed in Southern California and Los Angeles. The IGED enhanced genomic surveillance through predictive modeling and monitoring concerning evolutionary trends such as recombination and saltations in persistent infections. Development of the IGED highlighted the need for standardized data requirements, sustained funding for sequencing, incentives for data submission, and interdisciplinary collaboration to build an effective genomic surveillance system. This framework for linking genomic and epidemiologic data has not only generated critical insights for SARS‑CoV‑2 but also provided the foundation for CDPH and other public health organizations to develop similar IGED‑like systems for other priority pathogens as genomic surveillance expands.

Journal Article

Multi-omics Mendelian randomization integrating RNA-seq, eQTL and pQTL data revealed CPXM1 as a potential drug target for osteoporosis.

Osteoporosis, a prevalent skeletal disorder characterized by decreased bone mineral density and increased fracture risk, continues to be a major global health concern. Traditional treatments for osteoporosis have limited efficacy and safety profiles, highlighting the need for novel therapeutic targets. This study integrates multi-omics data, including RNA-seq, expression quantitative trait loci (eQTL), and protein quantitative trait loci (pQTL) data, through Mendelian randomization (MR) to identify potential drug targets for osteoporosis. By leveraging bidirectional two-sample MR analysis, we identified CPXM1 (Carboxypeptidase X, M14 family member 1) as a novel gene that is causally linked to osteoporosis risk. Through transcriptomic and proteomic validation, we demonstrate that CPXM1 was upregulated in aged bone tissues and osteoporotic conditions in both human and murine models. Gene set enrichment analysis (GSEA) revealed significant dysregulation of bone homeostasis pathways, including increased extracellular matrix degradation and suppression of osteoblast differentiation in aged mice. Furthermore, phenome-wide association studies (PheWAS) confirmed minimal off-target effects of CPXM1, reinforcing its potential as a therapeutic target. Finally, computational drug repurposing predicted several promising drug candidates, including Doxorubicin, 5-Fluorouracil, and 2-Methylcholine, which may target CPXM1 pathways for osteoporosis treatment. These findings highlight CPXM1 as a potential biomarker and therapeutic target, offering new avenues for osteoporosis therapy.

Osteoporosis

The gross national health product: a proposed population health index.

A population health status index designated as the gross national health product (GNHP) is proposed as a general measure of the health of nations or population groups. The GNHP integrates mortality and disability data into a single number in units of disability-free life years lived per 100,000 population. It is based primarily on mortality ratios and life expectancies of component age groups of the population, modified by their respective disability experiences. A computational example with data currently available on U.S. geographic regions from publications of the National Center for Health Statistics shows that the GNHP was highest in the West, indicating the highest number of disability-free years lived. Because of simplicity in its computation and interpretation, the GNHP can be used by health systems agencies (HSAs) in monitoring their performance or in conducting comparative studies.

Adolescent

Global Genomic Surveillance.

Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and environmental information within a One Health framework, while addressing critical challenges in governance, equity, and interoperability. The discussion covers the entire genomic surveillance workflow, from sample collection and sequencing to bioinformatic analysis and phylogenetic inference, and highlights the transformative role of artificial intelligence (AI) in predictive surveillance. By analyzing global initiatives, operational barriers, and emerging technologies, this chapter underscores the necessity of sustainable, equitable, and interoperable genomic systems to proactively address current and future infectious disease threats.

Humans

A doubly robust framework for addressing outcome-dependent selection bias in multi-cohort EHR studies.

Selection bias can hinder accurate estimation of association parameters in binary disease risk models using non-probability samples like electronic health records (EHRs). The issue is compounded when participants are recruited from multiple clinics/centers with varying selection mechanisms that may depend on the disease/outcome of interest. Traditional inverse-probability-weighted (IPW) methods, based on constructed parametric selection models, often struggle with misspecifications when selection mechanisms vary across cohorts. This paper introduces a new Joint Augmented Inverse Probability Weighted (JAIPW) method, which integrates individual-level data from multiple cohorts collected under potentially outcome-dependent selection mechanisms, with data from an external probability sample. JAIPW offers double robustness by incorporating a flexible auxiliary score model to address potential misspecifications in the selection models. We outline the asymptotic properties of the JAIPW estimator, and our simulations reveal that JAIPW achieves up to 6 times lower relative bias and 5 times lower root mean square error (RMSE) compared to the best performing joint IPW methods under scenarios with misspecified selection models. Applying JAIPW to the Michigan Genomics Initiative (MGI), a multi-clinic EHR-linked biobank, combined with external national probability samples, resulted in cancer-sex association estimates closely aligned with national benchmark estimates. We also analyzed the association between cancer and polygenic risk scores (PRS) in MGI to illustrate a situation where the exposure variable is not measured in the external probability sample.

Selection Bias

AI-Driven Precision Medicine in Alzheimer's Disease: Drug Repurposing, Digital Therapeutics and Clinical Decision Support.

Alzheimer's Disease (AD) is a neurodegenerative disease that causes significant clinical, social, and economic burden worldwide. Despite improvements in understanding its multifaceted pathogenesis, current treatments are mostly symptomatic and ineffective across varied patient populations. To overcome these constraints, AI-driven precision medicine allows tailored risk assessment, treatment selection, and disease monitoring. This review covers AI's role in AD precision medicine, focusing on drug repurposing, digital therapies and clinical decision support systems. Machine and deep learning models are used to predict medication response, integrate heterogeneous data sources such as genomics, transcriptomics, neuroimaging and electronic health records, and uncover pharmacogenomic treatment success factors. The paper covers AIenabled precision pharmacology, including tailored dosing algorithms, adaptive therapeutic monitoring, and adverse drug reaction prediction. Bioinformatics-based target identification, network pharmacology, graphbased AI models, virtual screening, and real-world and clinical data validation are emphasized in AI-driven medication repurposing. AI-powered digital treatments like personalized cognitive training platforms, wearable- derived digital biomarkers, virtual and mixed reality interventions, adherence monitoring, and digital twins for therapy optimization have been discussed. AI-based clinical decision support systems are also thoroughly assessed for clinical value, accuracy, and explainability in disease subtyping, trajectory prediction, and risk stratification in preclinical and prodromal AD. Despite these promises, data heterogeneity, algorithmic bias, legal barriers, and privacy concerns exist. Federated learning enables safe multi-center collaboration and hybrid AI-human approaches, and it represents the future. AI's ability to alter AD care opens the door to precision medicine paradigms that use repurposed medications, digital tools and intelligent decision-making to improve patient outcomes.

Alzheimer’s disease

Artificial Intelligence in Predicting Systemic Complications From Retinal Findings: A New Frontier in Precision Medicine.

Innovations in retinal imaging technologies and growing evidence from retinal imaging of systemic and neurodegenerative diseases have begun to explore the utility of retinal imaging in diagnosing these conditions. Since the retina shares embryological origins with the central nervous system and reflects systemic microvascular characteristics, it is well positioned for noninvasive observation of patients' systemic and neural health. Moreover, accessibility of retinal imaging has improved with the increasing number of ophthalmology clinics. Rapid improvements in various deep learning (DL) tools have also catalyzed the automation of retinal imaging analysis. Systems that utilize DL for retinal imaging are being developed to assist with disease recognition, clinical judgment, and prognostic assessment of systemic health. Various imaging modalities are being integrated with existing genomic and clinical data to estimate an individual's predisposition to certain conditions. Contrary to many existing reviews, the objective of this review is to synthesize the most recent clinical and technological evidence on DL-based diagnostic systems for retinal imaging, with a focus on how different network architectures and their combinations have been developed, validated, and applied across systemic disease detection and prediction. Specifically, this review examines the datasets, model validation approaches, and automated diagnostic systems reported in recent literature. It discusses the extent to which these advancements address existing barriers toward real-time diagnostic application across clinical disciplines. Integrating retinal imaging with DL is an innovative and promising approach to precision medicine and health risk reduction.

artificial intelligence

MULTIPREVENT: Integrated screening for smoking-related multimorbidity using low-dose chest computed tomography.

OBJECTIVES: Tobacco consumption, combined with individual genetic predispositions, contributes to an age-dependent risk not only for lung cancer but also for other non-communicable diseases (NCDs) such as cardiovascular disease (CVD), chronic obstructive pulmonary disease (COPD), osteoporosis, and diabetes. The MULTIPREVENT project aims to validate whether low-dose computed tomography (LDCT) of the chest, combined with simple biomarkers, functional tests, and genomic profiling, can serve as an effective tool for comprehensive health assessment and risk prediction of multimorbidity in adults. STUDY DESIGN: The study is based on a prospective epidemiological design involving 3000 participants from the MOLTEST-BIS lung cancer screening cohort (2016-2018). These participants, aged 50-79 years (during MOLTEST-BIS) and with a smoking history of at least 30 pack-years, will undergo two follow-up assessments in 2025-2027 and 2030-2032. METHODS: Each follow-up includes LDCT, spirometry, standardized blood pressure measurement, anthropometric evaluation, biomarker assessment (lipid profile, lipoprotein(a), glycated haemoglobin), and health-related questionnaires. Genetic profiling will be performed using the Illumina Infinium Global Screening Arrays approach to identify inherited predispositions to major NCDs. All data, clinical, imaging (including radiomics), molecular, and genetic, will be integrated through machine learning algorithms to develop AI-based risk prediction models. RESULTS: The MULTIPREVENT study is expected to generate a wide range of scientific, clinical, and infrastructural results that will serve as a foundation for future public health initiatives in integrated prevention. CONCLUSIONS: By linking imaging and biochemical markers, genetic susceptibility, and clinical parameters within a longitudinal design, MULTIPREVENT will establish data-driven, AI-supported prevention strategies aimed at reducing morbidity and mortality among adults exposed to tobacco. The project will also serve as a model for population-based multimorbidity prevention programs.

Humans

Clinical patient management and the integrated health information system.

Emerging progress in clinical applications of patient care computing is identified. The essential clinical skill is understanding what data are appropriate in any given patient care situation and extracting enough information to make the correct management decision. The value of the computer has less to do with the internal intellectual process of diagnosis than its contribution to the more manifest actions in support of clinical patient management. Techniques with which the computer is assisting in improving the clinical decisionmaking process are reviewed, and a mechanism to link them to active patient care settings is described. In addition, a trend toward the integration of various independent subsystems, so that expensive resources can be optimized for patient needs, is noted.

Computers

Scalable, open-access and multidisciplinary data integration pipeline for climate-sensitive diseases.

Climate-sensitive infectious diseases pose an important challenge for human, animal and environmental health and it has been estimated that over half of known human pathogenic diseases can be aggravated by climate change. While climatic and weather conditions are important drivers of transmission of vector-borne diseases, socio-economic, behavioural, and land-use factors as well as the interactions among them impact transmission dynamics. Analysis of drivers of climate-sensitive diseases require rapid integration of interdisciplinary data to be jointly analysed with epidemiological (including genomic and clinical) data. Current tools for the integration of multiple data sources are often limited to one data type or rely on proprietary data and software. To address this gap, we develop a scalable and open-access pipeline for the integration of multiple spatio-temporal datasets that requires only the declaration of the country and temporal range and resolution of the study. The tool is locally deployable and can easily be integrated into existing climate-disease-modelling applications. We demonstrate the utility of the tool for dengue modelling in Vietnam where epidemiological data are legally required to remain local. We include a pipeline for bias correction of climate data to enhance their quality for downstream modelling tasks. The Dengue Advanced Readiness Tools-Pipeline empowers users by simplifying complex download, correction, and aggregation steps, fostering data-driven discovery of relationships between infectious diseases and their drivers in space and time, and enhancing reproducibility in research. Additional modules and datasets can be added to the existing ones to make the pipeline extendable to use cases other than the ones presented here.

automated workflows

Beyond data and technology: the need for new thinking to enable the era of precision prevention.

BACKGROUND: Global flagship initiatives increasingly advocate for proactive health maintenance to alleviate the growing burden on reactive, disease-focused healthcare systems. Precision prevention is conceived as the targeted modulation of causal pathways across the disease continuum, from latent risk and pre-disease states to clinical manifestation, surpassing conventional public health prevention strategies that prioritise managing population-level risk factors. Traditional discovery and implementation models, however, remain poorly aligned with the pace and breadth of scientific and technological advances. This review outlines key barriers to scaling precision prevention and argues for the integration of conceptual, methodological, and policy perspectives into a single implementation‑oriented framework. MAIN: Individualised risk stratification lies at the core of precision prevention. Genomics serves as a stable substrate for lifetime susceptibility assessment, while meaningful prediction in multifactorial chronic disease requires additional risk monitoring using dynamic intermediate molecular markers and high-resolution exposomic data. Machine learning and other artificial intelligence (AI) methods are increasingly helpful tools for integrating large, heterogeneous and temporally structured real-world data to generate personalised predictions of health trajectories. Trustworthy AI-enabled risk prediction or decision-support systems are expected to provide transparency about model logic, assumptions and performance. In discovery, existing diagnostic classifications and conventional case-control designs can obscure mechanistic heterogeneity. Shifting toward precision phenotyping and biologically grounded disease redefinition could reveal a new layer of molecular understanding. Evidence generation strategies that reflect the temporal change of disease, including high‑risk enrichment, surrogate endpoints, and adaptive, trajectory-based monitoring, are particularly important for common conditions with prolonged latency periods (e.g., cancer, cardiovascular disease). Features often dismissed as "noise", such as stochastic molecular variation and minimal exposures, may in fact encode meaningful individual-level signals and thus merit investigation. CONCLUSION: To shift healthcare from reactive treatment toward proactive health maintenance requires coordinated action from stakeholders to reshape the pillars of discovery, reform outcome assessments and modernise implementation strategies.

Humans

Enterocutaneous Fistula-Associated Sepsis and Mortality: Development and Validation of a Multimodal Artificial Intelligence Prediction Model.

BACKGROUND: Predicting enterocutaneous fistula (ECF)-associated sepsis and mortality poses significant challenges in digital health care due to the disease's complexity and heterogeneous clinical manifestations. Current approaches that rely on single-modal data or traditional scoring systems often fail to capture the intricate immune-inflammatory dynamics and multisystem involvement in patients with ECF. OBJECTIVE: This study aims to develop an artificial intelligence (AI)-driven multimodal fusion model integrating clinical, imaging, and transcriptomic data for early prediction of ECF-associated sepsis and 28-day mortality, addressing the limitations of conventional single-dimensional models. METHODS: This study leveraged publicly available datasets (Medical Information Mart for Intensive Care III [MIMIC-III], electronic Intensive Care Unit [eICU], and The Cancer Genome Atlas) to construct a multimodal framework. Clinical parameters were processed using Extreme Gradient Boosting, abdominal imaging features were extracted via convolutional neural networks, and transcriptomic profiles were analyzed with variational autoencoders. A Transformer-based fusion network was employed for joint prediction and validated through cross-validation and external testing. Key features were identified using Shapley Additive Explanations and Local Interpretable Model-Agnostic Explanations interpretability algorithms, while immune regulatory mechanisms were explored via weighted gene co-expression network analysis. RESULTS: The multimodal model achieved an area under the curve (AUC) of 0.89 for predicting sepsis and 28-day mortality, outperforming unimodal models (clinical-only model, AUC 0.72, and imaging-only model, AUC 0.78). Critical predictors included Sequential Organ Failure Assessment score, lactate levels, intra-abdominal free fluid on imaging, and immunoregulatory genes (programmed death-ligand 1 [PD-L1] and indoleamine 2,3-dioxygenase 1 [IDO1]). Mechanistic analysis revealed distinct immune reprogramming in patients with sepsis, characterized by increased regulatory T cells and M2 macrophages, along with downregulated cluster of differentiation 8+ (CD8+) T cells. CONCLUSIONS: This multimodal AI model offers an innovative digital solution in medical informatics, enabling precise early risk stratification for ECF-associated sepsis. By integrating multisource data and providing interpretable insights into immune-inflammatory pathways, the model enhances health care quality for patients with ECF and paves the way for personalized intervention strategies.

Humans

Understanding Suicide through Coroners' Narratives: implications for primary care from a mixed‑methods study of 157 Coroners' reports.

Suicide is a major public health concern, and general practice is often a recent point of contact before death. While mental illness is well recognised, the broader social and contextual factors influencing suicide risk remain under-reported in primary care and epidemiological research Aim To describe the demographic, clinical, and psychosocial characteristics of individuals who died by suicide, integrating coronial quantitative data with qualitative narrative accounts to identify implications for primary/ secondary care and public health. Design and setting Explanatory sequential mixed‑methods study of 157 consecutive deaths by suicide recorded by coroners (2018-19) across five English local authorities. Method Demographic, clinical, and social data were extracted from coroners' records and summarised descriptively. Narrative case summaries were coded and analysed thematically to identify contextual, relational, and service factors preceding death. Results Of 157 individuals: 79% were male; 65% lived in the most deprived IMD quintile; 85% had a diagnosed mental health condition; 62% had a long‑term physical illness; 41% had a previous suicide attempt. About half consulted a GP in the preceding three months; mental health featured in about half of those consultations. Common stressors were relationship breakdown (37.2%), housing instability (22.1%), and work pressures (18.2%). Seven interlinked themes were identified: Mental health; Alcohol/Substance use, Physical health; Social connectedness; Life course trauma, Socioeconomic and Structural Vulnerability; Healthcare access. Service transitions were key vulnerability points Conclusion Coroners' records offer important insights into the complex circumstances preceding suicide and highlight opportunities for GPs to recognise intersectional complexity and support integrated, cross-sector suicide prevention approaches.

General Practice

Patterns of antimicrobial resistance genes in pathogens across One Health sectors in Ireland: an in silico approach.

As part of a rapid risk assessment, an in silico approach was used to detect antimicrobial resistance (AMR) in pathogenic isolates from humans, animals, and the environment. A total of 11,670 genomic data sets were retrieved from the NCBI Pathogen Detection system for Ireland, which represented 47 pathogenic species, including Salmonella enterica, Escherichia coli/Shigella spp., Staphylococcus aureus, Klebsiella pneumoniae, and Enterococcus faecium. Identifying the most critical pathogenic strains over time is essential, as these organisms significantly contribute to mortality, morbidity, and hospitalization. The analysis identified 799 antimicrobial resistance genes (ARGs), including their allelic diversity, 117 plasmid replicons, and 274 virulence factors. Several critical ARGs, particularly those conferring resistance to beta-lactams, aminoglycosides, quinolones, and colistin, were common across isolates originating from human, animal, and environmental sources, suggesting shared resistance profiles across One Health sectors. Klebsiella pneumoniae, E. coli/Shigella spp., S. enterica, and S. aureus were the dominant hosts of these ARGs and associated mobile genetic elements. Increasing resistance across major antibiotic classes aligned with trends reported across other European countries. This study provides a national-scale in silico comparison of AMR across pathogens and One Health sectors using publicly available genomic data. The findings help reinforce Ireland's AMR surveillance by showing which resistance genes are present and how they spread across critical pathogens in humans, animals, and the environment. These findings highlight the urgent need for improved antibiotic stewardship and integrated One Health surveillance to limit the emergence and spread of AMR.IMPORTANCEAntimicrobial resistance (AMR) is a growing threat to human, animal, and environmental health. This study used publicly available genomic data to identify antimicrobial resistance genes (ARGs) in key bacterial pathogens circulating in Ireland. By analyzing over 11,000 genomes from humans, animals, and the environment, we found that several dangerous resistance genes, including those against last-resort antibiotics, were widespread across different sources. The study highlights which bacteria and resistance genes are most critical and how they may spread between humans, animals, and the environment. These insights provide a national snapshot of AMR, supporting more effective monitoring and prevention strategies. By revealing patterns of resistance and modes of transmission, our findings underscore the importance of coordinated antibiotic stewardship and One Health approaches to slow the emergence and spread of resistant infections, protecting public health and ensuring antibiotics remain effective.

Humans

KG-Microbe: Building modular and scalable knowledge graphs for microbiome and microbial sciences.

BACKGROUND: The integration of many disparate forms of data is essential for understanding the microbial world and its interaction with the environment and human health. Doing so is particularly challenging in the context of microbe-host and microbe-microbe interactions that contribute to health or environmental outcomes. There are thousands of relevant microbial species, and millions of interactions among those microbes and with their environment or host. Integrated information (e.g., about host and microbial physiology, genetics, and metabolism) facilitates deeper understanding of complex mechanisms and helps interpret correlative results. RESULTS: The KG-Microbe construction framework is a novel approach to harmonizing bacterial and archaeal data in the form of a findable, accessible, interoperable, reusable and AI-ready knowledge graph (KG). Starting from a core KG with organismal traits, environments, and growth preferences and the integration of established ontologies, the framework generates a hierarchy of related KGs targeting specific use cases, including the human microbiome in the context of disease, or environmental microbiomes. The framework supports customizable taxa subsets representing communities or clades of interest. Evaluations of the KG-Microbe KGs through a series of competency questions demonstrate the accuracy and effectiveness of the data harmonization, and the utility of the resulting KGs in studies of inflammatory bowel disease and Parkinson's disease. Finally, the predictive and environmental capabilities of the KGs are demonstrated by predicting growth preferences using graph features. CONCLUSIONS: The KG-Microbe framework unifies microbial contexts in a single resource to support integrative analyses across biomedical, host, and environmental domains. KG-Microbe is a flexible, modular enabling technology for humans and machine learning methods to uncover candidate mechanistic explanations of microbial associations.

Microbiota

Identification of rare maternal copy number variants by genome-wide analysis of noninvasive prenatal screening data in 113,017 pregnant women.

OBJECTIVES: Knowledge of copy number variants (CNVs) is relevant to maternal and fetal health and can be obtained from noninvasive prenatal screening (NIPS) of pregnancy. However, genome-wide analysis of maternal CNVs using NIPS data has not been conducted in large populations. METHODS: For CNV analysis, the human genome was segmented into 10 kilobase pairs (Kb) bins, and the relative sequencing depth of each bin was calculated. The circular binary segmentation algorithm was used to estimate CNVs. Detected CNVs from two pregnancies of the same participant were compared to validate the reproducibility. All CNVs were merged into CNV regions (CNVRs) to evaluate their frequency, distributions, and relationship with disease-related genes and regions. RESULTS: In this study, 113,017 pregnant women were recruited. A total of 363,886 CNVs larger than 50 Kb were detected in 101,779 individuals and merged into 43,005 CNVRs. For evaluating the reproducibility of CNVs, 90.18% of deletions and 88.07% of duplications were consistent. In general, 78.13% of individuals carried CNVRs that overlapped protein-coding genes, while 14.76% overlapped OMIM genes. We detected 246 novel CNVRs, 134 (54.47%) involving protein-coding genes. For the perspective of maternal-fetal health, we identified 4,984 (4.41%) individuals as carriers of 5,243 CNVs containing known pathogenic or likely pathogenic regions, including 22q11.2 region and DMD gene.. CONCLUSIONS: NIPS sequencing data is a reliable source for maternal CNV detection. These CNVs constitute an integrate component in maternal-fetal health management.

Humans