PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Experimental validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Combinatorial microRNA target predictions.

MicroRNAs are small noncoding RNAs that recognize and bind to partially complementary sites in the 3' untranslated regions of target genes in animals and, by unknown mechanisms, regulate protein production of the target transcript. Different combinations of microRNAs are expressed in different cell types and may coordinately regulate cell-specific target genes. Here, we present PicTar, a computational method for identifying common targets of microRNAs. Statistical tests using genome-wide alignments of eight vertebrate genomes, PicTar's ability to specifically recover published microRNA targets, and experimental validation of seven predicted targets suggest that PicTar has an excellent success rate in predicting targets for single microRNAs and for combinations of microRNAs. We find that vertebrate microRNAs target, on average, roughly 200 transcripts each. Furthermore, our results suggest widespread coordinate control executed by microRNAs. In particular, we experimentally validate common regulation of Mtpn by miR-375, miR-124 and let-7b and thus provide evidence for coordinate microRNA control in mammals.

Algorithms↗

Genomic exploration and in silico prioritization of putative COX-2-targeting metabolites from Streptomyces sp. VITGV156 (MCC 4965).

INTRODUCTION: Streptomyces species represent an important source of bioactive natural products, yet systematic genome-guided prioritization of metabolites targeting cyclooxygenase-2 (COX-2/PTGS2) remains limited. This study aimed to investigate the biosynthetic potential of Streptomyces sp. VITGV156 (MCC 4965) using an integrated genome mining and computational drug discovery pipeline. METHODS: Whole-genome sequencing, functional annotation, antiSMASH v7.0.1-based biosynthetic gene cluster (BGC) prediction, LC-MS/MS metabolomic profiling, SwissADME analysis, target prediction, disease association mapping, molecular docking against PTGS2 (PDB: 5IKR), and PASS bioactivity prediction were performed to prioritize putative bioactive metabolites. RESULTS: Genome analysis identified 29 predicted biosynthetic gene clusters, including clusters associated with geosmin, ectoine, albaflavenone, hopene, coelichelin, and SapB, together with several cryptic clusters exhibiting low similarity to known pathways. LC-MS/MS metabolomic profiling provided experimental support for active secondary metabolite production under the cultivation conditions employed. Computational prioritization identified PTGS2 (COX-2) as a biologically relevant target. Molecular docking demonstrated favorable binding affinities and interaction profiles for several predicted metabolites within the PTGS2 catalytic pocket. PASS analysis further suggested potential anticancer-related biological activities that require experimental validation. DISCUSSION: These findings demonstrate the utility of integrating genome mining, metabolomic profiling, and computational drug discovery for prioritizing natural-product candidates. Streptomyces sp. VITGV156 (MCC 4965) represents a promising source of biosynthetic diversity and provides a genome-guided framework for identifying putative COX-2-targeting natural products for future experimental validation rather than confirming metabolite production or biological activity.

COX-2 (PTGS2)↗

Dosimetric characterization of radioactive sources employed in prostate cancer therapy.

PURPOSE: Three types of radiation sources are employed currently in the radiation treatment of prostate cancer, namely, external, implant, and high-dose-rate (HDR) sources using an afterloader method. The present article provides a detailed dosimetric characterization of several commercially available implant sources and an HDR source employing the same stochastic code and dataset. METHODS AND MATERIALS: The radioactive implants considered are (125)I seeds: models 6701, 6702 and 6711, (103)Pd seed: model 200, and a high-dose-rate (192)Ir source: microSelectron-HDR model V7.0x. Detailed modeling of the sources and their associated X-rays and gamma rays has been carried out using the stochastic code MCNP4C. A sensitivity study has been conducted to quantify effects of varying the composition and density of the tissue equivalent material, and a dosimetric comparison is made for different media (tissue equivalent, solid-water, water, and air). Furthermore, a set of measurements using thermoluminescent dosimeters has been done to provide experimental validation of some of the calculational results obtained. RESULTS: Effectively, high-precision dosimetric values (Monte-Carlo statistical 1-sigma error <1%) are provided in tabulated form over a wide range to enable therapy planning as well as to check numerical values calculated by other methods. A subset of calculated dosimetric values has been experimentally validated by using thermoluminescent dosimeters. CONCLUSIONS: A detailed comparison of results obtained for the radial dose distribution function, anisotropy factor, and dose rate constant as defined in the TG-43 protocol has indicated reasonable agreement with the values reported in the literature.

Brachytherapy↗

Potential for enlarging DNA memory: the validity of experimental operations of scaled-up nested primer molecular memory.

DNA is an attractive memory unit because of its immense information density. Here, we describe a memory model made of DNA, called Nested Primer Molecular Memory (NPMM). NPMM consists of many DNA strands, and each DNA strand consists of two areas: a data area and a data address area. When the address of target data is specified, only the target data can be extracted from NPMM. In this paper, we evaluate the validity of the basic operations of NPMM and then discuss the feasibility of scaled-up NPMM through some laboratory experiments. In the latter, we deal with scaled-up NPMM simulated by the Concentration Scaling method.

Algorithms↗

Integrative genomics based identification of potential human hepatocarcinogenesis-associated cell cycle regulators: RHAMM as an example.

DNA microarray has been widely used to examine gene expression profile of different human tumors. The information generated from microarray analysis usually represents the overall range of cancer-associated abnormality associated with gene regulation. In order to identify key regulatory genes involved in carcinogenesis of human cancer, hypothesis driven data mining of the microarray data plus experimental validation becomes a critical approach in the post-genome era. Here, we present an integrative genomic analysis of published microarray data and homolog gene database. Over 20,000 genes were examined to reveal 16 genes specific to vertebrates, cell cycle G2/M regulated, and overexpressed in human HCC. Using Affymetrix microarray analysis, we found that all 16 genes were up-regulated in human HCC. Among these 16 genes, we experimentally validated the up-regulation of receptor for hyaluronan-mediated motility (RHAMM) in different cell model systems. We first confirmed elevation of RHAMM in the G2/M phase of synchronized HeLa cells. We also found that RHAMM had an elevated level of expression in all the HCC samples we examined and it was induced during the G2/M phase of regenerating mouse hepatocytes after partial hepatectomy. Thus, the expression of RHAMM appears to be tightly regulated during mammalian cell cycle G2/M progression. The ectopic overexpression of RHAMM in 293T cells resulted in the accumulation of cells at G2/M phase. RHAMM-induced mitotic arrest of cells was predominantly in the prophase. Taken together, using an integrated functional genomic approach, we have uncovered a set of genes that may play specific roles in cell cycle progression and in HCC development. To elucidate the function of these genes in cell cycle regulation may shed light on the control mechanism of human HCC in the future.

Amino Acid Sequence↗

A framework for computational and experimental methods: identifying dimerization residues in CCR chemokine receptors.

Solving relevant biological problems requires answering complex questions. Addressing such questions traditionally implied the design of time-consuming experimental procedures which most of the time are not accessible to average-sized laboratories. The current trend is to move towards a multidisciplinary approach integrating both theoretical knowledge and experimental work. This combination creates a powerful tool for shedding light on biological problems. To illustrate this concept, we present here a descriptive example of where computational methods were shown to be a key aspect in detecting crucial players in an important biological problem: the dimerization of chemokine receptors. Using evolutionary based sequence analysis in combination with structural predictions two CCR5 residues were selected as important for dimerization and further validated experimentally. The experimental validation of computational procedures demonstrated here provides a wealth of valuable information not obtainable by any of the individual approaches alone.

Amino Acid Sequence↗

C-peptide as a measure of the secretion and hepatic extraction of insulin. Pitfalls and limitations.

The large and variable hepatic extraction of insulin is a major obstacle to our ability to quantitate insulin secretion accurately in human subjects. The evidence that C-peptide is secreted from the beta cell in equimolar concentration with insulin, but not extracted by the liver to any significant degree, has provided a firm scientific basis for the use of peripheral C-peptide concentrations as a semiquantitative marker of beta cell secretory activity in a variety of clinical situations. Thus, plasma C-peptide has proved to be extremely valuable in the study of the natural history of type 1 diabetes, to monitor insulin secretion in patients with insulin antibodies, and as an adjunct in the investigation of patients with hypoglycemic disorders. The use of the peripheral C-peptide concentration to accurately quantitate the rate of insulin secretion is more controversial. This is mainly because understanding of the kinetics and metabolism of C-peptide under different conditions is incomplete. Unfortunately, sufficient quantities of human C-peptide are not available to allow the experimental validation of the mathematical formulae that have been proposed for the calculation of insulin secretion from peripheral C-peptide concentrations. Until it is possible to perform such experiments, the accuracy of studies that have derived insulin secretion rates from peripheral C-peptide levels will remain uncertain. The assumption that the peripheral C-peptide:insulin molar ratio can be used as a reflection of hepatic insulin extraction has not been experimentally validated. The marked difference in the plasma half-lives of insulin and C-peptide complicates the interpretation of changes in their ratios.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult↗

Localization of cerebral arterovenous malformations using digital angiography.

Since 1989 we performed stereotactic radiotherapy treatments of cerebral arterovenous malformations (AVM), estimating three-dimensional (3-D) localization and shape of target volumes by the Leksell stereotactic helmet on two orthogonal radiographic projections. Due to the limitations of this method, we developed a new technique for the localization of the target volume using digital subtraction angiography (DSA) and digital image processing. To achieve this result we first developed a method to correct nonlinear distortion of DSA images using spatial relocation of image pixels based on a calibration grid. We then developed an algorithm for localization of the target volume using two independent DSA projections. Target volume coordinates in the helmet system are calculated using two DSA acquisitions taken with a free angle (approximately 90 degrees), one in the AP and the other in the LL direction. The helmet can be freely positioned between the x-ray source and the image plane. The projections of eight reference points inserted in the helmet at a known location, are used to calculate the transformation matrix between the two coordinate systems. We performed numerical and experimental validation of the system. A hypothetical random error (up to 2 mm) on image coordinates of the reference points allowed to determine that the error in target localization was less than 0.2 mm. Using DSA images of target points with a known location within a phantom, the error between calculated and actual location was, on average, 0.30+/-0.13 mm (mean+/-SD), with a maximum error of 0.49 mm. The results of numerical and experimental validations show that the system we have developed allows fast and accurate localization of the center of the target volume and it is suitable for efficient guiding during stereotactic radiosurgery of AVM.

Algorithms↗

Theoretical model of restriction endonuclease HpaI in complex with DNA, predicted by fold recognition and validated by site-directed mutagenesis.

Type II restriction enzymes are commercially important deoxyribonucleases and very attractive targets for protein engineering of new specificities. At the same time they are a very challenging test bed for protein structure prediction methods. Typically, enzymes that recognize different sequences show little or no amino acid sequence similarity to each other and to other proteins. Based on crystallographic analyses that revealed the same PD-(D/E)XK fold for more than a dozen case studies, they were nevertheless considered to be related until the combination of bioinformatics and mutational analyses has demonstrated that some of these proteins belong to other, unrelated folds PLD, HNH, and GIY-YIG. As a part of a large-scale project aiming at identification of a three-dimensional fold for all type II REases with known sequences (currently approximately 1000 proteins), we carried out preliminary structure prediction and selected candidates for experimental validation. Here, we present the analysis of HpaI REase, an ORFan with no detectable homologs, for which we detected a structural template by protein fold recognition, constructed a model using the FRankenstein monster approach and identified a number of residues important for the DNA binding and catalysis. These predictions were confirmed by site-directed mutagenesis and in vitro analysis of the mutant proteins. The experimentally validated model of HpaI will serve as a low-resolution structural platform for evolutionary considerations in the subgroup of blunt-cutting REases with different specificities. The research protocol developed in the course of this work represents a streamlined version of the previously used techniques and can be used in a high-throughput fashion to build and validate models for other enzymes, especially ORFans that exhibit no sequence similarity to any other protein in the database.

Amino Acid Sequence↗

Molecular characterization of a gene encoding N-myristoyl transferase (NMT) from Triticum aestivum (bread wheat).

Myristoyl-CoA:protein N-myristoyl transferase (NMT; EC 2.3.1.97) acylates the Gly residue abutting the N-terminal Met with a myristic acid following the removal of the Met residue in certain eukaryotic proteins, and in some cases myristoylation is essential to cell growth and survival. We report the cloning of a full-length cDNA encoding NMT from Triticum aestivum (TaNMT). The cDNA included a predicted open reading frame of 1317 nucleotides, which encoded a predicted protein of 438 amino acids containing all of the residues that are important for NMT activity. The TaNMT amino acid and nucleotide sequences were compared with NMTs from 14 other species encompassing a wide array of taxonomic groups. Among the experimentally validated NMTs, TaNMT was most similar to that of Arabidopsis thaliana. Southern blot analysis of wheat genomic DNA showed that TaNMT is encoded by a single copy gene, with one copy per haploid genome. We expressed TaNMT in Escherichia coli cells and determined that the recombinant protein possessed NMT activity, catalyzing the N-myristoylation of peptides from known or putatively myristoylated proteins from plants and animals without a strong preference for the plant peptides. TaNMT is the second experimentally validated plant NMT sequence and the first from a monocotyledonous species.

Acylation↗

Rehabilitation of executive functioning: an experimental-clinical validation of goal management training.

Two studies assessed the effects of a training procedure (Goal Management Training, GMT), derived from Duncan's theory of goal neglect, on disorganized behavior following TBI. In Study 1, patients with traumatic brain injury (TBI) were randomly assigned to brief trials of GMT or motor skills training. GMT, but not motor skills training, was associated with significant gains on everyday paper-and-pencil tasks designed to mimic tasks that are problematic for patients with goal neglect. In Study 2, GMT was applied in a postencephalitic patient seeking to improve her meal-preparation abilities. Both naturalistic observation and self-report measures revealed improved meal preparation performance following GMT. These studies provide both experimental and clinical support for the efficacy of GMT toward the treatment of executive functioning deficits that compromise independence in patients with brain damage.

Adult↗

Influence of Major Histocompatibility Complex (MHC) Diversity on Immune Modulation, Pathogenesis, and Control of Lumpy Skin Disease Virus.

INTRODUCTION: Lumpy Skin Disease Virus (LSDV), a member of the genus Capripoxvirus within the family Poxviridae, is an economically important transboundary viral pathogen affecting cattle and water buffalo. The disease causes severe production losses through decreased milk yield, infertility, hide damage, reduced growth performance, and occasional mortality. The rapid geographic spread of LSDV, together with its vectorborne transmission and emerging recombinant strains, has intensified the need for improved understanding of viral pathogenesis, host immune responses, and effective prevention strategies. In particular, the role of the bovine Major Histocompatibility Complex (BoLA/MHC) in regulating antiviral immunity, disease susceptibility, and vaccine responsiveness has gained increasing scientific attention. METHODS: This review summarises the published literature related to the epidemiology, transmission, structure, pathogenesis, diagnosis, prevention, and control of LSDV, with special emphasis on the immunological and molecular role of bovine MHC molecules. Relevant studies concerning BoLA-mediated antigen presentation, immunoinformaticsbased epitope prediction, vaccine development, antiviral drug repurposing, molecular docking, genomic surveillance, and diagnostic approaches, including PCR- and ELISAbased assays, were critically evaluated. Recent advances in computational biology, molecular virology, and host-pathogen interaction studies were also reviewed. RESULTS: The reviewed studies demonstrate that Lumpy Skin Disease Virus (LSDV) possesses a complex double-stranded DNA genome enabling immune modulation and efficient transmission through arthropod vectors such as mosquitoes, ticks, and biting flies. Disease progression involves systemic viral replication, vascular injury, dermal necrosis, and inflammatory skin lesions. Real-time PCR remains the most sensitive diagnostic method for early detection, while ELISA supports surveillance. Evidence highlights the central role of bovine Major Histocompatibility Complex (BoLA) molecules in antigen presentation and T-cell activation. Computational studies identified promising BoLA-binding epitopes and repurposed antiviral candidates, including ivermectin, theaflavin, canagliflozin, and tepotinib, for future therapeutic development. DISCUSSION: Current evidence indicates that effective LSDV control requires integration of molecular diagnostics, vector management, vaccination, and host immunogenetics. BoLAguided immunoinformatics provides promising opportunities for developing multi-epitope vaccines, although experimental validation remains essential. Similarly, repurposed antiviral candidates require comprehensive in vivo and pharmacological evaluation before clinical application. Future research should focus on elucidating viral immune-evasion mechanisms, validating predicted epitopes, and translating computational findings into practical vaccines and therapeutics for sustainable disease control. CONCLUSION: Lumpy Skin Disease continues to pose a major threat to global cattle health and livestock economies. Advances in molecular diagnostics, genomic surveillance, antiviral drug discovery, and BoLA-guided vaccine design provide promising opportunities for improved disease control. Understanding the interaction between LSDV and the bovine MHC system is essential for developing next-generation vaccines, immunotherapeutics, and precision disease-management strategies. Future research should prioritise experimental validation of predicted epitopes, large-scale vaccine trials, and mechanistic studies on host-virus immune interactions to establish effective and sustainable global control programs for LSDV.

BoLA↗

A sequence-based classifier distinguishes phenotype-associated genes from other gene models in plants.

Only a small fraction of annotated plant genes possess experimentally validated associations with specific phenotypes. Phenotype-associated genes have distinct structural, molecular, and evolutionary characteristics compared with nonvalidated gene models. Here, we develop a simple classifier that uses sequence and evolutionary features, which can be generated for any species with an annotated reference genome assembly, to accurately distinguish phenotype-associated genes from both the overall population of annotated gene models and a specific set of genes identified as being tolerant of premature stop mutations. A model trained solely on genes from maize (Zea mays) identifies and prioritizes rice (Oryza sativa) and Arabidopsis (Arabidopsis thaliana) genes that are highly enriched in genes with experimentally validated links to phenotypes in both of these evolutionarily distant species. Gene models predicted to have a higher probability of being linked to phenotypes display patterns consistent with known biological properties of phenotype-associated genes. Notably, the sets of genes predicted to have a high probability of being linked to phenotype variation do not consist exclusively of well-characterized gene families but included many uncharacterized gene families carrying domains of unknown function. The quantitative scores generated by this model offer a valuable resource for prioritizing and exploring the vast number of uncharacterized gene models in plants, reducing the risk of failure in future reverse genetic efforts and potentially accelerating gene discovery and functional annotation in crops.

Phenotype↗

Alcohol and the validation of experimental aggression paradigms: the Taylor reaction time procedure.

The aim was to find out whether intoxicated and sober subjects would calibrate a shock scale to the same objective level and whether shocks received would be subjectively experienced in the same way in terms of pain and discomfort. The intention was also to replicate previous studies attributing an aggression enhancing effect to alcohol. The subjective ratings were made within the Taylor reaction time 'aggression machine'. Twenty-four subjects were randomly assigned to either an alcohol or a control group. The former drank 0.8 ml of pure alcohol/kg body wt. Results indicated no differences among groups on a shock setting measure of aggression under unprovoked or provoked conditions and no differences in level of calibration of shocks or in subjective ratings of pain and discomfort. These results were contrary to all predictions and are discussed as indirectly supportive of an hypothesis stating that this version of the 'aggression machine' may not generate a valid measure of aggression when used in alcohol research.

Adult↗

Establishing a statistic model for recognition of steroid hormone response elements.

Identification of hormone response elements (HREs) is essential for understanding the mechanism of hormone-regulated gene expression. To date, there has been a lack of effective bioinformatics tools for recognition of specific HRE such as Progesterone Response Elements (PRE). In this paper, a comprehensive survey and comparison of in silico methods is conducted for establishing a more accurate statistic model. Homogeneity of steroid HRE is analyzed and a reliable training dataset is constructed through extensive searching for experimentally validated response elements from more than 150 literature sources. Based on the observation that the verified HREs carry di-nucleotide preservation in comparison with uniform nucleotide distributions, both mono and di-nucleotide Position Weight Matrices are computed to extract the statistic pattern of the positions. It is followed by the sequence transition pattern recognition using a specifically designed profile Hidden Markov Model. Reciprocal combination of the statistic and transition patterns significantly improves the performance of the model in terms of higher sensitivity and specificity. Upon acquisition of the putative response elements in the promoter areas of vertebrate genes, a qualitative scheme is applied to assess the probability for each gene to be a hormone primary target. Using >650 records of experimentally validated steroid hormone response elements, a high sensitivity level of 73% and high specificity level of one prediction per 8.24 kb is reached, allowing this model to be used for further prediction of primary target genes through the analysis of their upstream promoters, for human or other vertebrate genomes of interest. Additional documents, supplementary data and the web-based program developed for response elements prediction are freely available for academic research at . Submission of putative gene promoter regions for recognition of potential regulatory PREs can be as long as 5 kb.

Animals↗

Minimizing immunogenicity of the SDR-grafted humanized antibody CC49 by genetic manipulation of the framework residues.

The murine mAb CC49 specifically recognizes a tumor-associated glycoprotein (TAG)-72, which is expressed on the majority of human carcinomas. This Ab has potential applications in the diagnosis and treatment of human carcinomas. However, patients receiving murine CC49 generate human anti-murine Ab (HAMA) responses, preventing repeated administration of the Ab for effective treatment. To minimize the HAMA response, two versions of humanized CC49 (HuCC49) were developed: (a) HuCC49 and (b) HuCC49V10 (V10). HuCC49 was developed by grafting the CC49 CDRs, while V10 was generated by grafting only the specificity determining residues (SDRs) of the CC49 onto the frameworks of the human Abs. During the generation of both HuCC49 and V10, a few murine framework residues that were believed to be essential for the integrity of the Ag-binding site were retained. However, the indispensability of these residues for the Ag-binding activity of CC49 has not been experimentally validated. In this study, an array of V10 variants were generated by replacing, by site-specific mutagenesis, the murine framework residues that were retained in the humanized Ab with their counterparts in the human templates. The variants were tested for their (a) Ag-binding activity and (b) reactivity to sera from patients who were previously administered murine CC49 in a clinical trial. One such variant, V59, compared to the parental V10, shows a significant decrease in its reactivity to the anti-variable region Abs present in the patients' sera, while it binds to the TAG-72 Ag with a slightly higher affinity. Variant 59, which is expected to be minimally immunogenic because of its low sera reactivity, is a potentially useful clinical reagent against human carcinomas. In this study, we show for the first time that experimental validation rather than reliance on the protein data bank (PDB) should be the criterion for the indispensability of framework residues for the humanization of any murine Ab to retain its Ag-binding property and reduce its immunogenicity in patients.

Animals↗

High-throughput identification, database storage and analysis of SNPs in EST sequences.

Single nucleotide polymorphisms (SNPs) are the most frequent form of DNA variation and disease-causing mutations in many genes. Due to their abundance and slow mutation rate within generations, they are thought to be the next generation of genetic markers that can be used in a myriad of important biological, genetic, pharmacological, and medical applications. There are several strategies both experimental, and in-silico for SNP discovery and mapping. Experimental SNP discovery consists of a number of labourious steps that make this process complex and expensive. In-silico discovery has been proposed as an alternative discovery method that makes use and takes advantage of large data sets with potential SNP information that have been generated with other purposes and have not been used as a SNP information source yet. However, in order to successfully apply the in-silico method to large data sets, the following challenges need to be addressed: First it is necessary to build an integrated SNP pipeline that handles data processing steps smoothly from the beginning (collecting sequence information) to end (SNPs in the database). Also, SNP detection tool parameters have to be optimized to satisfy specific goals of the project. Finally, SNP data could not be fully used until the in-silico method is validated experimentally. In this paper we present a design and implementation of an in-silico SNP detection software pipeline that exploits the existence of large EST (expressed sequence tag) data sets and effectively addresses the above challenges. First, the pipeline allows for smooth data transition between its different components by implementing data interfaces that translate the data formats of the different tools in the different stages. Second, we optimized PolyBayes parameters for SNP detection in maize EST. Finally, we implemented a user interface that along with the database structure created allows the scientist to perform preliminary analysis of the data and to perform basic statistics on the SNP data prior to experimental validation. The pipeline works with two different types of sequence assemblers (PHRAP (http://www.phrap.org/) and CAT from DoubleTwist (http://www.doubletwist.com/). It uses a Bayesian engine for SNP detection (PolyBayes), selects relevant polymorphism information which is then uploaded into a database. We detected 2439 SNPs and 822 insertion deletions (INDELs) with a PolyBayes probability higher than 0.99 on the public set of 68,000 maize ESTs. The user interface allowed us analyzing the polymorphism information right after discovery in several ways that allowed us to gain insight into the distribution and significance of the newly acquired data.

Animals↗

Multi-level aggregation analysis of microbiome composition and host gene expression reveals associations with systemic and local immunity.

The human gut microbiome plays a critical role in immune regulation, yet the molecular links between microbiome composition and host gene expression remain incompletely understood. We analyzed associations between host gene expression and microbiome composition in a cohort of 315 healthy individuals, integrating microarray-based gene expression data from three intestinal sites (ileum, transverse colon, and rectum) and six immune cell types with microbiome sequencing data. Using a hierarchical feature aggregation strategy combining principal component analysis, clustering, and covariate correction, we discovered significant associations primarily related to immunity. While microbial profiles were similar across the three intestinal sites, the transverse colon yielded the most "microbiome-host gene expression" associations. Among the immune cell types, CD8+ cells showed the highest number of associations. The first principal component of microbiome composition, reflecting a gradient from commensals (e.g., Ruminococcaceae and Christensenellaceae) to proinflammatory taxa ([Ruminococcus] gnavus and Lachnoclostridium), correlated with the expression of TNF-&#x3b1;-linked genes (HMOX1, CPI17, HSD3B2, and SLC5A1). Among individual genera, Catenibacterium abundance was associated with gene expression in both intestinal and immune cells, including negative associations with MRPS21 (related to mitochondrial function) in the transverse colon and with CD8+ gene programs related to T cell differentiation. These findings align with emerging evidence implicating mitochondrial dysfunction in intestinal inflammation. Our results identify multi-level associations between the gut microbiome and host gene expression, suggesting potential mechanisms by which microbiota shape local and systemic immunity and vice versa. The implicated genes and taxa represent candidates for experimental validation to improve understanding of host-microbiome homeostasis and its disruption in disease.IMPORTANCEThe gut microbiome and immune system are engaged in a complex interplay throughout human life. While most associative studies focus on case-control comparisons-typically examining patients with conditions such as inflammatory bowel disease or metabolic diseases-less is known about the molecular links between the microbiome and immune system in healthy individuals. In this study of a large cohort of healthy individuals, we addressed this gap by applying multiscale modeling to tackle the high dimensionality of host-microbiome data. We identified multi-level associations between microbiome composition and host gene expression in both intestinal tissues and immune cells. These findings offer a valuable reference for understanding baseline host-microbiome communication and highlight molecular candidates-such as TNF-&#x3b1;-related genes and mitochondrial pathways-for future experimental validation.

Humans↗