PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Network inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Low-order conditional independence graphs for inferring genetic networks.

As a powerful tool for analyzing full conditional (in-)dependencies between random variables, graphical models have become increasingly popular to infer genetic networks based on gene expression data. However, full (unconstrained) conditional relationships between random variables can be only estimated accurately if the number of observations is relatively large in comparison to the number of variables, which is usually not fulfilled for high-throughput genomic data. Recently, simplified graphical modeling approaches have been proposed to determine dependencies between gene expression profiles. For sparse graphical models such as genetic networks, it is assumed that the zero- and first-order conditional independencies still reflect reasonably well the full conditional independence structure between variables. Moreover, low-order conditional independencies have the advantage that they can be accurately estimated even when having only a small number of observations. Therefore, using only zero- and first-order conditional dependencies to infer the complete graphical model can be very useful. Here, we analyze the statistical and probabilistic properties of these low-order conditional independence graphs (called 0-1 graphs). We find that for faithful graphical models, the 0-1 graph contains at least all edges of the full conditional independence graph (concentration graph). For simple structures such as Markov trees, the 0-1 graph even coincides with the concentration graph. Furthermore, we present some asymptotic results and we demonstrate in a simulation study that despite their simplicity, 0-1 graphs are generally good estimators of sparse graphical models. Finally, the biological relevance of some applications is summarized.

Algorithms↗

Algebraic comparison of metabolic networks, phylogenetic inference, and metabolic innovation.

BACKGROUND: Comparison of metabolic networks is typically performed based on the organisms' enzyme contents. This approach disregards functional replacements as well as orthologies that are misannotated. Direct comparison of the structure of metabolic networks can circumvent these problems. RESULTS: Metabolic networks are naturally represented as directed hypergraphs in such a way that metabolites are nodes and enzyme-catalyzed reactions form (hyper)edges. The familiar operations from set algebra (union, intersection, and difference) form a natural basis for both the pairwise comparison of networks and identification of distinct metabolic features of a set of algorithms. We report here on an implementation of this approach and its application to the procaryotes. CONCLUSION: We demonstrate that metabolic networks contain valuable phylogenetic information by comparing phylogenies obtained from network comparisons with 16S RNA phylogenies. The algebraic approach to metabolic networks is suitable to study metabolic innovations in two sets of organisms, free living microbes and Pyrococci, as well as obligate intracellular pathogens.

Algorithms↗

Inference of S-system models of genetic networks using a cooperative coevolutionary algorithm.

MOTIVATION: To resolve the high-dimensionality of the genetic network inference problem in the S-system model, a problem decomposition strategy has been proposed. While this strategy certainly shows promise, it cannot provide a model readily applicable to the computational simulation of the genetic network when the given time-series data contain measurement noise. This is a significant limitation of the problem decomposition, given that our analysis and understanding of the genetic network depend on the computational simulation. RESULTS: We propose a new method for inferring S-system models of large-scale genetic networks. The proposed method is based on the problem decomposition strategy and a cooperative coevolutionary algorithm. As the subproblems divided by the problem decomposition strategy are solved simultaneously using the cooperative coevolutionary algorithm, the proposed method can be used to infer any S-system model ready for computational simulation. To verify the effectiveness of the proposed method, we apply it to two artificial genetic network inference problems. Finally, the proposed method is used to analyze the actual DNA microarray data.

Algorithms↗

Inference network-based analyses of the histopathological effects of androgen deprivation on prostate cancer.

The evaluation of prostate cancer histology following hormonal therapy often represents a diagnostic problem for the pathologist. Previous studies have shown that an inference or Bayesian belief network (BBN) offers a descriptive classifier useful for the accurate analysis of morphological changes in individual cases of prostate neoplasia. Three different BBNs were evaluated in 94 cancer foci present in 20 radical prostatectomy (RP) specimens and in the matching biopsies in which the initial diagnosis of prostatic adenocarcinoma was made. Ten RP specimens were from patients treated with total androgen ablation or combination endocrine therapy (CET) before surgery. The first and second BBN allowed the identification with high certainty of the cancer foci present in the biopsies and RP specimens, as well as their Gleason grade, the belief value often being close to 1.0. The results of the second BBN showed a good correspondence between the Gleason grade given in the biopsies and that in the RP specimens, except in the surgical material of the treated patients, in which upgrading was always present. The third BBN showed the existence of three subgroups in treated RP specimens, one with morphological effect, another with poor effect, and the third with the histology of untreated (i.e. unaffected) cancer. In conclusion, an inference network-based analysis allows the characterization of treated prostate cancers according to the degree of histopathological change.

Adenocarcinoma↗

Automatically inferred Markov network models for classification of chromosomal band pattern structures.

A structural pattern recognition approach to the analysis and classification of metaphase chromosome band patterns is presented. An operational method of representing band pattern profiles as sharp edged idealized profiles is outlined. These profiles are nonlinearly scaled to a few, but fixed number of "density" levels. Previous experience has shown that profiles of six levels are appropriate and that the differences between successive bands in these profiles are suitable for classification. String representations, which focuses on the sequences of transitions between local band pattern levels, are derived from such "difference profiles." A method of syntactic analysis of the band transition sequences by dynamic programming for optimal (maximal probability) string-to-network alignments is described. It develops automatic data-driven inference of band pattern models (Markov networks) per class, and uses these models for classification. The method does not use centromere information, but assumes the p-q-orientation of the band pattern profiles to be known a priori. It is experimentally established that the method can build Markov network models, which, when used for classification, show a recognition rate of about 92% on test data. The experiments used 200 samples (chromosome profiles) for each of the 22 autosome chromosome types and are designed to also investigate various classifier design problems. It is found that the use of a priori knowledge of Denver Group assignment only improved classification by 1 or 2%. A scheme for typewise normalization of the class relationship measures prove useful, partly through improvements on average results and partly through a more evenly distributed error pattern. The choice of reference of the p-q-orientation of the band patterns is found to be unimportant, and results of timing of the execution time of the analysis show that recent and efficient implementations can process one cell in less than 1 min on current standard hardware. A measure of divergence between data sets and Markov network models is shown to provide usable estimates of experimental classification performance.

Chromosome Banding↗

Diagnostic decision support for prostate lesions.

The diagnostic evaluation of premalignant and malignant lesions of the prostate may benefit from the application of an inference network. Used as a diagnostic decision support system, an inference network provides standardized assessment of diagnostic clues which is supported by computer graphics and comparison imagery, uncertainty management by possibility and probabilistic schemes and the systematic combination of different pieces of diagnostic evidence. This assessment results in a numeric measure of belief in the final diagnosis.

Diagnosis, Computer-Assisted↗

Inferring gene regulatory networks from time series data using the minimum description length principle.

MOTIVATION: A central question in reverse engineering of genetic networks consists in determining the dependencies and regulating relationships among genes. This paper addresses the problem of inferring genetic regulatory networks from time-series gene-expression profiles. By adopting a probabilistic modeling framework compatible with the family of models represented by dynamic Bayesian networks and probabilistic Boolean networks, this paper proposes a network inference algorithm to recover not only the direct gene connectivity but also the regulating orientations. RESULTS: Based on the minimum description length principle, a novel network inference algorithm is proposed that greatly shrinks the search space for graphical solutions and achieves a good trade-off between modeling complexity and data fitting. Simulation results show that the algorithm achieves good performance in the case of synthetic networks. Compared with existing state-of-the-art results in the literature, the proposed algorithm exceptionally excels in efficiency, accuracy, robustness and scalability. Given a time-series dataset for Drosophila melanogaster, the paper proposes a genetic regulatory network involved in Drosophila's muscle development. AVAILABILITY: Available from the authors upon request.

Algorithms↗

Elucidation of directionality for co-expressed genes: predicting intra-operon termination sites.

MOTIVATION: In this paper, we present a novel framework for inferring regulatory and sequence-level information from gene co-expression networks. The key idea of our methodology is the systematic integration of network inference and network topological analysis approaches for uncovering biological insights. RESULTS: We determine the gene co-expression network of Bacillus subtilis using Affymetrix GeneChip time-series data and show how the inferred network topology can be linked to sequence-level information hard-wired in the organism's genome. We propose a systematic way for determining the correlation threshold at which two genes are assessed to be co-expressed using the clustering coefficient and we expand the scope of the gene co-expression network by proposing the slope ratio metric as a means for incorporating directionality on the edges. We show through specific examples for B. subtilis that by incorporating expression level information in addition to the temporal expression patterns, we can uncover sequence-level biological insights. In particular, we are able to identify a number of cases where (1) the co-expressed genes are part of a single transcriptional unit or operon and (2) the inferred directionality arises due to the presence of intra-operon transcription termination sites. AVAILABILITY: The software will be provided on request. SUPPLEMENTARY INFORMATION: http://www.phys.psu.edu/~ralbert/pdf/gma_bioinf_supp.pdf

Algorithms↗

Inferring cellular networks using probabilistic graphical models.

High-throughput genome-wide molecular assays, which probe cellular networks from different perspectives, have become central to molecular biology. Probabilistic graphical models are useful for extracting meaningful biological insights from the resulting data sets. These models provide a concise representation of complex cellular networks by composing simpler submodels. Procedures based on well-understood principles for inferring such models from data facilitate a model-based methodology for analysis and discovery. This methodology and its capabilities are illustrated by several recent applications to gene expression data.

Bayes Theorem↗

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis↗

Median-joining networks for inferring intraspecific phylogenies.

Reconstructing phylogenies from intraspecific data (such as human mitochondrial DNA variation) is often a challenging task because of large sample sizes and small genetic distances between individuals. The resulting multitude of plausible trees is best expressed by a network which displays alternative potential evolutionary paths in the form of cycles. We present a method ("median joining" [MJ]) for constructing networks from recombination-free population data that combines features of Kruskal's algorithm for finding minimum spanning trees by favoring short connections, and Farris's maximum-parsimony (MP) heuristic algorithm, which sequentially adds new vertices called "median vectors", except that our MJ method does not resolve ties. The MJ method is hence closely related to the earlier approach of Foulds, Hendy, and Penny for estimating MP trees but can be adjusted to the level of homoplasy by setting a parameter epsilon. Unlike our earlier reduced median (RM) network method, MJ is applicable to multistate characters (e.g., amino acid sequences). An additional feature is the speed of the implemented algorithm: a sample of 800 worldwide mtDNA hypervariable segment I sequences requires less than 3 h on a Pentium 120 PC. The MJ method is demonstrated on a Tibetan mitochondrial DNA RFLP data set.

Algorithms↗

Co-mutation Based Genetic Networks to Infer Temporal Mutation Dynamics in Ancient Human Mitochondrial Genomes.

The evolutionary history of Homo sapiens is marked by complex interactions between environmental, cultural, and genetic factors. To investigate the molecular signatures of these processes, we analyzed ancient mitochondrial DNA (mtDNA) across temporal and geographic contexts using principles of co-occurrence of minor alleles defined as co-mutation, through spatiotemporal co-mutation networks of variable sites. Haplogroup-based assessments of variable sites revealed a major transition from foraging to agrarian lifestyles during the Copper-Bronze Age. Genetic network analyses demonstrated that COX and CYB loci exhibited distinct temporal dynamics, with their interactions modulated by NADH dehydrogenase genes in a geological age-dependent manner. To complement the network approach, we constructed phylogeny-based gene interaction networks and assessed polymorphism-to-divergence from chimpanzee ratios. The tree-based networks displayed topologies consistent with co-mutation analyses but showed reduced gene-gene connectivity. Polymorphism/divergence analysis further indicated that the CYB gene has been under long-term purifying selection, whereas ATP6, COX, and NADH dehydrogenase genes experienced episodic purifying selection aligned with distinct historical phases. Collectively, our findings demonstrate that network-based analysis of ancient mtDNA provides insights into early human lifestyle transitions and haplogroup diversification, contributing to the evolutionary foundations of modern human populations.

Ancient humans↗

SimpleMicrobiome: An integrated web-based platform for streamlined microbiome data analysis and visualization.

Microbiome studies require multiple analytical steps after initial sequence processing. These steps commonly include data harmonization, preprocessing, taxonomic profiling, diversity analysis, differential abundance testing, predictive modeling, network inference, and preparation of publication-ready outputs. Although robust packages are available for many of these tasks, routine use often depends on command-line workflows, repeated data reformatting, and method-specific scripting. These requirements can limit accessibility for experimental researchers and complicate consistent analysis across interdisciplinary teams. We developed SimpleMicrobiome, a web-based R Shiny platform that integrates established microbiome analysis methods into a single interactive downstream workflow. The application accepts standard abundance, taxonomy, and metadata tables, supports interactive preprocessing and sample filtering, and provides modules for taxa profile visualization, alpha and beta diversity analysis, ANCOM-BC2 and MaAsLin2 differential abundance testing, Random Forest modeling with SHAP-based interpretation, microbial association network inference using SparCC and SPIEC-EASI through NetCoMi, correlation heatmaps, and dbRDA/CAP-style association biplots. The platform is implemented as a modular Shiny application so that preprocessing choices are propagated across downstream analyses, results can be exported as figures and tables, and the same application can be run through the public server, source-code installation, or a Docker image. SimpleMicrobiome consolidates major downstream microbiome analysis tasks in an accessible browser-based environment while retaining links to established analytical frameworks. The platform may reduce technical barriers for non-programming users, improve consistency across exploratory and reporting-oriented analyses, and support collaborative microbiome research. The public application is available at https://simplemicrobiome.mglab.org, the source code is available at https://github.com/yjcho2252/SimpleMicrobiome, and a Docker image for local deployment is available at https://hub.docker.com/r/mglab2252/simplemicrobiome.

differential abundance↗

ReGAIN: a bioinformatics platform for assessing probabilistic co-occurrence between resistance genes in bacterial pathogens.

MOTIVATION: Multidrug-resistant bacterial pathogens continue to rise globally, yet scalable methods are needed to infer how resistance determinants co-occur across pathogen populations and to quantify conditional dependencies underlying co-occurrence and shared genetic context. RESULTS: We present ReGAIN (Resistance Gene Association and Inference Network), an open-source platform that applies Bayesian network structure learning to infer probabilistic, conditional dependency relationships among antibiotic resistance, heavy metal tolerance, stress response, and virulence determinants in bacteria. In contrast to pairwise co-occurrence analyses, ReGAIN reports conditional probabilities, relative risks, and absolute risk differences with confidence intervals to prioritize candidate relationships for downstream prioritization. Applied across ESKAPEE pathogens, ReGAIN recapitulated established resistance gene relationships and identified additional candidate patterns consistent with co-selection and shared genetic context. Together, these results support scalable, reproducible population-wide analysis of resistance networks for surveillance, comparative genomics and epidemiology. AVAILABILITY: ReGAIN analyses are performed using Python v3.11.5 and R v4.4.1 and is available as open-source software through Bioconda at {https://anaconda.org/bioconda/regain-cli}. Source code and documentation can be found at {https://github.com/ERBringHorvath/regain_CLI}. All genomes used in this publication were downloaded from the National Center for Biotechnology Information database. Large supplementary tables and results data from the ESKAPEE pathogen example network analyses can be downloaded from https://figshare.com/articles/dataset/ReGAIN_command_line_software_and_supplemental_figures_/28959431.

Computational Biology↗

Functional transcriptomes: comparative analysis of biological pathways and processes in eukaryotes to infer genetic networks among transcripts.

Microarray technology enables us to monitor large changes in transcripts at any given time. The compilation of these data makes possible the comparison of such gene expression data on a genome-wide scale. As comparisons of genome sequence data yield new biological insights, comparative analyses of transcriptome data also promise new discoveries regarding metabolic pathways and cellular processes. The coordinated expression of genes shows that these genes physically interact with each other or are part of the same cascade. We have produced one of the largest expression profiles of adult mice and developmental tissues. These data, as well as the data on yeast from previous reports, were used to see whether coordinated expression (with high correlation coefficient) is closely coupled to the actual cascade on the pathway map.

Animals↗

Mathematical methods for inferring regulatory networks interactions: application to genetic regulation.

This paper deals with the problem of reconstruction of the intergenic interaction graph from the raw data of genetic co-expression coming with new technologies of bio-arrays (DMA-arrays, protein-arrays, etc.). These new imaging devices in general only give information about the asymptotical part (fixed configurations of co-expression or limit cycles of such configurations) of the dynamical evolution of the regulatory networks (genetic and/or proteic) underlying the functioning of living systems. Extracting the casual structure and interaction coefficients of a gene interaction network from the observed configurations is a complex problem. But if all the fixed configurations are supposedly observed and if they are factorizable into two or more subsets of values, then the interaction graph possesses as many connected components as the number of factors and the solution is obtained in polynomial time. This new result allows us for example to partly solve the topology of the genetic regulatory network ruling the flowering in Arabidopsis thaliana .

Algorithms↗

JCell--a Java-based framework for inferring regulatory networks from time series data.

MOTIVATION: JCell is a Java-based application for reconstructing gene regulatory networks from experimental data. The framework provides several algorithms to identify genetic and metabolic dependencies based on experimental data conjoint with mathematical models to describe and simulate regulatory systems. Owing to the modular structure, researchers can easily implement new methods. JCell is a pure Java application with additional scripting capabilities and thus widely usable, e.g. on parallel or cluster computers. AVAILABILITY: The software is freely available for download at http://www-ra.informatik.uni-tuebingen.de/software/JCell.

Algorithms↗

Inferring phylogenetic networks by the maximum parsimony criterion: a case study.

Horizontal gene transfer (HGT) may result in genes whose evolutionary histories disagree with each other, as well as with the species tree. In this case, reconciling the species and gene trees results in a network of relationships, known as the "phylogenetic network" of the set of species. A phylogenetic network that incorporates HGT consists of an underlying species tree that captures vertical inheritance and a set of edges which model the "horizontal" transfer of genetic material. In a series of papers, Nakhleh and colleagues have recently formulated a maximum parsimony (MP) criterion for phylogenetic networks, provided an array of computationally efficient algorithms and heuristics for computing it, and demonstrated its plausibility on simulated data. In this article, we study the performance and robustness of this criterion on biological data. Our findings indicate that MP is very promising when its application is extended to the domain of phylogenetic network reconstruction and HGT detection. In all cases we investigated, the MP criterion detected the correct number of HGT events required to map the evolutionary history of a gene data set onto the species phylogeny. Furthermore, our results indicate that the criterion is robust with respect to both incomplete taxon sampling and the use of different site substitution matrices. Finally, our results show that the MP criterion is very promising in detecting HGT in chimeric genes, whose evolutionary histories are a mix of vertical and horizontal evolution. Besides the performance analysis of MP, our findings offer new insights into the evolution of 4 biological data sets and new possible explanations of HGT scenarios in their evolutionary history.

Algorithms↗