PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Network inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Use of artificial neural networks to evaluate the effectiveness of riverbank filtration.

Riverbank filtration (RBF) is a low-cost water treatment technology in which surface water contaminants are removed or degraded as the infiltrating water moves from the river/lake to the pumping wells. The removal or degradation of contaminants is a combination of physicochemical and biological processes. This paper illustrates the development and application of three types of artificial neural networks (ANNs) to estimate the effectiveness of two RBF facilities in the US. The feed-forward back-propagation network (BPN) and radial basis function network (RBFN) model prediction results produced excellent agreement with measured data at a correlation coefficient above 0.99 for filtrate water quality parameters, including temperature as well as turbidity, heterotrophic bacteria, and coliform removal. In comparison, the fuzzy inference system network (FISN) predicted only temperature and bacteria removal with reasonable accuracy. It is shown that the predictive performances of the ANNs depend on the model structure and model inputs.

Bacteria↗

A gene coexpression network for bovine skeletal muscle inferred from microarray data.

We present the application of large-scale multivariate mixed-model equations to the joint analysis of nine gene expression experiments in beef cattle muscle and fat tissues with a total of 147 hybridizations, and we explore 47 experimental conditions or treatments. Using a correlation-based method, we constructed a gene network for 822 genes. Modules of muscle structural proteins and enzymes, extracellular matrix, fat metabolism, and protein synthesis were clearly evident. Detailed analysis of the network identified groupings of proteins on the basis of physical association. For example, expression of three components of the z-disk, MYOZ1, TCAP, and PDLIM3, was significantly correlated. In contrast, expression of these z-disk proteins was not highly correlated with the expression of a cluster of thick (myosins) and thin (actin and tropomyosins) filament proteins or of titin, the third major filament system. However, expression of titin was itself not significantly correlated with the cluster of thick and thin filament proteins and enzymes. Correlation in expression of many fast-twitch muscle structural proteins and enzymes was observed, but slow-twitch-specific proteins were not correlated with the fast-twitch proteins or with each other. In addition, a number of significant associations between genes and transcription factors were also identified. Our results not only recapitulate the known biology of muscle but have also started to reveal some of the underlying associations between and within the structural components of skeletal muscle.

Adipose Tissue↗

Elucidation of gene interaction networks through time-lagged correlation analysis of transcriptional data.

The photosynthetic cyanobacterium Synechocystis sp. strain PCC 6803 uses a complex genetic program to control its physiological response to alternating light conditions. To study this regulatory program time-series experiments were conducted by exposing Synechocystis sp. to serial perturbations in light intensity. In each experiment whole-genome DNA microarrays were used to monitor gene transcription in 20-min intervals over 8- and 16-h periods. The data was analyzed using time-lagged correlation analysis, which identifies genetic interaction networks by constructing correlations between time-shifted transcription profiles with different levels of statistical confidence. These networks allow inference of putative cause-effect relationships among the organism's genes. Using light intensity as our initial input signal, we identified six groups of genes whose time-lagged profiles possessed significant correlation, or anti-correlation, with the light intensity. We expanded this network by using the average profile from each group of genes as a seed, and searching for other genes whose time-lagged profiles possessed significant correlation, or anti-correlation, with the group's average profile. The final network comprised 50 different groups containing 259 genes. Several of these gene groups possess known light-stimulated gene clusters, such as Synechocystis sp. photosystems I and II and carbon dioxide fixation pathways, while others represent novel findings in this work.

Cyanobacteria↗

On the functional equivalence of fuzzy inference systems and spline-based networks.

The conditions under which spline-based networks are functionally equivalent to the Takagi-Sugeno-model of fuzzy inference are formally established. We consider a generalized form of basis function network whose basis functions are splines. The result admits a wide range of fuzzy membership functions which are commonly encountered in fuzzy systems design. We use the theoretical background of functional equivalence to develop a hybrid fuzzy-spline net for inverse dynamic modeling of a hydraulically driven robot manipulator.

Fuzzy Logic↗

Inference of differential kinase interaction networks with KINference.

MOTIVATION: Differential kinase interaction networks (DKINs) are networks containing kinase-substrate links that are differentially active between two conditions. Existing methods are either able to predict condition-agnostic kinase-substrate links or condition-specific differential kinase activity, but do not provide differential kinase-substrate links. Moreover, existing methods for predicting kinase-substrate links usually rely on curated biochemical knowledge. Thus, there is a lack of data-driven DKIN inference methods that are also applicable when prior knowledge is scarce. RESULTS: To address this need, we present KINference. KINference combines computation of a baseline KIN representing the space of all possible kinase-substrate links with filters applied to nodes and edges to identify differentially active subnetworks that are relevant in the context of a specific phosphoproteomics dataset. For the node filters, we rely on functional relevance and differential phosphorylation scores; for the edge filters, we make use of prize-collecting Steiner trees and correlations between phosphorylation sites of kinases and their target proteins. Tests on two phosphoproteomics datasets (kinase inhibition in breast cancer cells, SARS-CoV-2 infection in Calu-3 cells) show that the proposed filters produce significant results in terms of overlap with known interactions between kinases and phosphorylation sites. Furthermore, a case study on the SARS-CoV-2 infection data, suggests a potential host pathway linked to virus replication, showcasing the process of hypothesis generation utilizing DKINs computed by KINference. AVAILABILITY AND IMPLEMENTATION: KINference is available as an R package at https://github.com/bionetslab/KINference and https://doi.org/10.5281/zenodo.15411150. Scripts to reproduce the results are available at https://github.com/bionetslab/KINference-Evaluation-Scripts and https://doi.org/10.5281/zenodo.15424599.

Humans↗

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics↗

Improved success of phenotype prediction of the human immunodeficiency virus type 1 from envelope variable loop 3 sequence using neural networks.

We have assembled two sets of HIV-1 V3 sequences with defined epidemiologic relationships associated with experimentally determined coreceptor usage or MT-2 cell tropism. These data sets were used for three purposes. First, they were employed to test existing methods for predicting coreceptor usage and MT-2 cell tropism. Of these methods, the presence of one basic amino acid at position 11 or 25 proved to be most reliable for both phenotypic classifications, although its predictive power for the X4 phenotype was less than 50%. Second, we used the sequence sets to train neural networks to infer coreceptor usage from V3 genotype with better success than the best available motif-based method, and with a predictive power equal to that of the best motif-based method for MT-2 cell tropism. Third, we used the sequence sets to reexamine patterns of variability associated with the different phenotypes, and we showed that the phenotype-associated sequence patterns could be reproduced from large sets of V3 sequences using phenotypes predicted by the trained neural network.

Algorithms↗

Chain functions and scoring functions in genetic networks.

One of the grand challenges of system biology is to reconstruct the network of regulatory control among genes and proteins. High throughput data, particularly from expression experiments, may gradually make this possible in the future. Here we address two key ingredients in any such 'reverse engineering' effort: The choice of a biologically relevant, yet restricted, set of potential regulation functions, and the appropriate score to evaluate candidate regulatory relations. We propose a set of regulation functions which we call chain functions, and argue for their ubiquity in biological networks. We analyze their complexity and show that their number is exponentially smaller than all boolean functions of the same dimension. We define two new scores: one evaluating the fitness of a candidate set of regulators of a particular gene, and the other evaluating a candidate function. Both scores use established statistical methods. Finally, we test our methods on experimental gene expression data from the yeast galactose pathway. We show the utility of using chain functions and the improved inference using our scores in comparison to several extant scores. We demonstrate that the combined use of the two scores gives an extra advantage. We expect both chain functions and the new scores to be helpful in future attempts to infer regulatory networks.

Galactose↗

Stable evolutionary signal in a yeast protein interaction network.

BACKGROUND: The recently emerged protein interaction network paradigm can provide novel and important insights into the innerworkings of a cell. Yet, the heavy burden of both false positive and false negative protein-protein interaction data casts doubt on the broader usefulness of these interaction sets. Approaches focusing on one-protein-at-a-time have been powerfully employed to demonstrate the high degree of conservation of proteins participating in numerous interactions; here, we expand his 'node' focused paradigm to investigate the relative persistence of 'link' based evolutionary signals in a protein interaction network of S. cerevisiae and point out the value of this relatively untapped source of information. RESULTS: The trend for highly connected proteins to be preferably conserved in evolution is stable, even in the context of tremendous noise in the underlying protein interactions as well as in the assignment of orthology among five higher eukaryotes. We find that local clustering around interactions correlates with preferred evolutionary conservation of the participating proteins; furthermore the correlation between high local clustering and evolutionary conservation is accompanied by a stable elevated degree of coexpression of the interacting proteins. We use this conserved interaction data, combined with P. falciparum/Yeast orthologs, as proof-of-principle that high-order network topology can be used comparatively to deduce local network structure in non-model organisms. CONCLUSION: High local clustering is a criterion for the reliability of an interaction and coincides with preferred evolutionary conservation and significant coexpression. These strong and stable correlations indicate that evolutionary units go beyond a single protein to include the interactions among them. In particular, the stability of these signals in the face of extreme noise suggests that empirical protein interaction data can be integrated with orthologous clustering around these protein interactions to reliably infer local network structures in non-model organisms.

Evolution, Molecular↗

A recurrent neural network for modelling dynamical systems.

We introduce a recurrent network architecture for modelling a general class of dynamical systems. The network is intended for modelling real-world processes in which empirical measurements of the external and state variables are obtained at discrete time points. The model can learn from multiple temporal patterns, which may evolve on different timescales and be sampled at non-uniform time intervals. We demonstrate the application of the model to a synthetic problem in which target data are only provided at the final time step. Despite the sparseness of the training data, the network is able not only to make good predictions at the final time step for temporal processes unseen in training, but also to reproduce the sequence of the state variables at earlier times. Moreover, we show how the network can infer the existence and role of state variables for which no target information is provided. The ability of the model to cope with sparse data is likely to be useful in a number of applications, including, in particular, the modelling of metal forging.

Models, Theoretical↗

A mammalian promoter model links cis elements to genetic networks.

An accurate identification of gene promoters remains an important challenge. Computational approaches for this problem rely on promoter sequence attributes that are believed to be critical for transcription initiation. Here we report a probabilistic model that captures two important properties of promoters, not used by previous methods, viz., the location preference and co-occurrence of promoter elements. Additionally, we found that many of the position-specific DNA elements are strongly linked with the function of the gene product. For instance, a highly conserved motif CCTTT at -1 position is strongly associated with protein synthesis, cellular and tissue development. Our comparative analysis of promoter classes reveals that the promoters devoid of CpG islands are more conserved and have fewer alternative transcription start sites. The discovered links between promoter elements and gene function allows us to infer genetic networks from promoter elements. The web server for the PSPA promoter predictor is available at /PSPA.

Animals↗

Direct measurement of the area expansion and shear moduli of the human red blood cell membrane skeleton.

The area expansion and the shear moduli of the free spectrin skeleton, freshly extracted from the membrane of a human red blood cell (RBC), are measured by using optical tweezers micromanipulation. An RBC is trapped by three silica beads bound to its membrane. After extraction, the skeleton is deformed by applying calibrated forces to the beads. The area expansion modulus K(C) and shear modulus mu(C) of the two-dimensional spectrin network are inferred from the deformations measured as functions of the applied stress. In low hypotonic buffer (25 mOsm/kg), one finds K(C) = 4.8 +/- 2.7 microN/m, mu(C) = 2.4 +/- 0.7 microN/m, and K(C)/mu(C) = 1.9 +/- 1.0. In isotonic buffer, one measures higher values for K(C), mu(C), and K(C)/mu(C), partly because the skeleton collapses in a high-ionic-strength environment. Some data concerning the time evolution of the mechanical properties of the skeleton after extraction and the influence of ATP are also reported. In the Discussion, it is shown that the measured values are consistent with estimates deduced from experiments carried out on the intact membrane and agree with theoretical and numerical predictions concerning two-dimensional networks of entropic springs.

Adenosine Triphosphate↗

STRING: a database of predicted functional associations between proteins.

Functional links between proteins can often be inferred from genomic associations between the genes that encode them: groups of genes that are required for the same function tend to show similar species coverage, are often located in close proximity on the genome (in prokaryotes), and tend to be involved in gene-fusion events. The database STRING is a precomputed global resource for the exploration and analysis of these associations. Since the three types of evidence differ conceptually, and the number of predicted interactions is very large, it is essential to be able to assess and compare the significance of individual predictions. Thus, STRING contains a unique scoring-framework based on benchmarks of the different types of associations against a common reference set, integrated in a single confidence score per prediction. The graphical representation of the network of inferred, weighted protein interactions provides a high-level view of functional linkage, facilitating the analysis of modularity in biological processes. STRING is updated continuously, and currently contains 261 033 orthologs in 89 fully sequenced genomes. The database predicts functional interactions at an expected level of accuracy of at least 80% for more than half of the genes; it is online at http://www.bork.embl-heidelberg.de/STRING/.

Algorithms↗

Integrated multi-omics profiling identifies aging-related molecular signatures and convergent interferon signaling in systemic lupus erythematosus.

BACKGROUND: Systemic lupus erythematosus (SLE) is characterized by chronic immune activation and molecular alterations that overlap with aging-related biological processes. However, how these alterations are organized across molecular layers and whether they converge on shared regulatory networks remain incompletely understood. METHODS: We performed an integrative multi-omics analysis combining in-house proteomic and phosphoproteomic data from 130 patients with SLE and 90 healthy controls (HCs) and publicly available transcriptomic datasets comprising 1,461 SLE patients. Proteins and phosphorylation sites were annotated using established aging-related gene resources. Differential protein abundance and phosphorylation changes were analyzed across disease-status and disease-activity comparisons. Nominal P-value thresholds were used for exploratory feature selection, whereas FDR-adjusted P values were used to assess robustness after multiple-testing correction. Kinase-substrate enrichment, transcription factor annotation, and cell-type-resolved transcriptomic comparison were used to explore potential regulatory programs. RESULTS: We identified 128 nominally altered proteins annotated to aging-related biological processes, including genomic instability, mitochondrial dysfunction, and epigenetic alterations. Phosphoproteomic analysis revealed 36 nominally altered phosphorylation sites, including previously unreported sites in IFI16 (S153, S780) and PKCδ (S507, S664). Clustering analysis demonstrated heterogeneous protein co-regulation patterns across disease states. Kinase activity inference suggested altered activity of TBK1 and IKKβ. TF analysis further highlighted STAT1, RELA, and PML as potential central nodes within the inferred regulatory network. Notably, these multi-omic alterations were not randomly distributed but showed convergence toward shared signaling pathways, particularly those related to interferon responses. CONCLUSIONS: This integrative multi-omics study identifies inflammatory and interferon-dominated molecular alterations in SLE PBMCs that overlap with aging-related biological processes and converge on shared regulatory networks. These findings provide a hypothesis-generating framework for investigating the intersection between chronic immune activation and aging-related molecular remodeling in SLE.

Humans↗

Comparative phylogenomics and transcriptional regulatory networks of AQPs, HSPs, and LEA proteins in salt-stressed Portulaca oleracea.

Soil salinization severely threatens global food security, necessitating systematic investigations of halophytes like Portulaca oleracea to decode the molecular mechanisms of environmental resilience. Utilizing an integrated framework of deep learning-based genome annotation (58,817 predicted genes; 96.5% BUSCO completeness), multi-tissue RNA-Seq, phylogenomics, and gene regulatory network (GRN) inference, the synergistic orchestration of 78 aquaporins (AQPs), 525 heat shock proteins (HSPs), and 119 late embryogenesis abundant (LEA) proteins was elucidated. The active transcriptome, encompassing 39,065 expressed loci, revealed a systemic growth-defense trade-off. Tissues displayed distinct adaptive mechanisms: leaves modulated intracellular water balance via specialized AQPs, whereas adult roots maintained proteostasis through robust HSP20/HSP70 induction. Phylogenomic clustering across 154 species demonstrated that salinity tolerance constitutes an evolutionary mosaic, identifying 81 halophyte-exclusive orthogroups and 1129 species-specific clusters. Comparative topology across six independent GRNs (4.2M-5.3 M edges) unmasked a highly modular transcriptional reprogramming strategy governed by a core apparatus of 22 stress-exclusive regulators, with functional enrichment heavily prioritizing protein dimerization and chromatin remodeling. Theoretically, the distinct convergence of Trihelix transcription factors with guard cell differentiation pathways offers a candidate transcriptomic framework to explain the plant's characteristic C4-CAM photosynthetic plasticity under severe osmotic pressure. Practically, these evolutionary blueprints and specific master switches transcend single-gene transgenic limitations. Utilizing these root-sustained and stress-inducible targets under localized promoters provides a naturally optimized, network-level precision engineering roadmap to transfer robust, compartmentalized halotolerance to sensitive glycophytic crops.

Gene Regulatory Networks↗

scPlantLLM: A Foundation Model for Exploring Single-cell Expression Atlases in Plants.

Single-cell RNA sequencing (scRNA-seq) provides unprecedented insights into plant cellular diversity by enabling high-resolution analyses of gene expression at the single-cell level. However, the complexity of scRNA-seq data, including challenges in batch integration, cell type annotation, and gene regulatory network (GRN) inference, demands advanced computational approaches. To address these challenges, we developed scPlantLLM, a Transformer model trained on millions of plant single-cell data points. Using a sequential pretraining strategy incorporating masked language modeling and cell type annotation tasks, scPlantLLM generates robust and interpretable single-cell data embeddings. When applied to Arabidopsis thaliana datasets, scPlantLLM excels in clustering, cell type annotation, and batch integration, achieving an accuracy of up to 0.91 in zero-shot learning scenarios. Furthermore, the model demonstrates an ability to identify biologically meaningful GRNs and subtle cellular subtypes, showcasing its potential to advance plant biology research. Compared to traditional methods, scPlantLLM outperforms in key metrics such as adjusted rand index (ARI), normalized mutual information (NMI), and silhouette score (SIL), highlighting its superior clustering accuracy and biological relevance. scPlantLLM represents a foundation model for exploring plant single-cell expression atlases, offering unprecedented capabilities to resolve cellular heterogeneity and regulatory dynamics across diverse plant systems. The code used in this study is available at https://github.com/compbioNJU/scPlantLLM.

Single-Cell Analysis↗

Influence of growth medium, age in vitro and spontaneous bioelectric activity on the distribution of sensory ganglion-evoked activity in spinal cord explants.

The role of serum added to the culture medium and of spontaneous bioelectric activity in the development of sensory afferent connections was studied, employing fetal mouse spinal cord explants with attached dorsal root ganglia (DRG) as an in vitro model system. Afferent DRG terminals in the cord explants were localized on the basis of 'fixed-latency' DRG-evoked action potentials, which were anatomically verified in several experiments using horseradish peroxidase histology. In serum-supplemented medium (HSSM), but not in chemically defined medium (CDM), those DRG fibers which grew into the dorsal side of the cord terminated predominantly within the dorsal cord region, and remained there throughout the experimental period (18-33 days in vitro). In contrast, ventrally entering fibers terminated equally in both the dorsal and the ventral cord regions in young cultures (18-24 days in vitro) but were no longer observed after 27 days in vitro. Cultures grown in HSSM with the addition of xylocaine, in order to chronically suppress spontaneous bioelectric activity, essentially corresponded (at 25-32 days in vitro) to the picture seen in the control series at the same age. On the basis of polysynaptic DRG-evoked responses in the cord, developmental changes in local neuronal networks were inferred which resulted in less spread of DRG-evoked activity with age in HSSM, and more spread with age in CDM-grown cultures. It is concluded that for the formation of selective DRG connections in the spinal cord: (i) a serum-borne factor plays a role: and (ii) functional activity is not required.

Animals↗

A mathematical framework for inferring connectivity in probabilistic neuronal networks.

We describe an approach for determining causal connections among nodes of a probabilistic network even when many nodes remain unobservable. The unobservable nodes introduce ambiguity into the estimate of the causal structure. However, in some experimental contexts, such as those commonly used in neuroscience, this ambiguity is present even without unobservable nodes. The analysis is presented in terms of a point process model of a neuronal network, though the approach can be generalized to other contexts. The analysis depends on the existence of a model that captures the relationship between nodal activity and a set of measurable external variables. The mathematical framework is sufficiently general to allow a large class of such models. The results are modestly robust to deviations from model assumptions, though additional validation methods are needed to assess the success of the results.

Algorithms↗