PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Network inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Reconciling gene expression data with known genome-scale regulatory network structures.

The availability of genome-scale gene expression data sets has initiated the development of methods that use this data to infer transcriptional regulatory networks. Alternatively, such regulatory network structures can be reconstructed based on annotated genome information, well-curated databases, and primary research literature. As a first step toward reconciling the two approaches, we examine the consistency between known genome-wide regulatory network structures and extensive gene expression data collections in Escherichia coli and Saccharomyces cerevisiae. By decomposing the regulatory network into a set of basic network elements, we can compute the local consistency of each instance of a particular type of network element. We find that the consistency of network elements is influenced by both structural features of the network such as the number of regulators acting on a target gene and by the functional classes of the genes involved in a particular element. Taken together, the approach presented allows us to define regulatory network subcomponents with a high degree of consistency between the network structure and gene expression data. The results suggest that targeted gene expression profiling data can be used to refine and expand particular subcomponents of known regulatory networks that are sufficiently decoupled from the rest of the network.

Computational Biology↗

A likelihood approach to analysis of network data.

Biological, sociological, and technological network data are often analyzed by using simple summary statistics, such as the observed degree distribution, and nonparametric bootstrap procedures to provide an adequate null distribution for testing hypotheses about the network. In this article we present a full-likelihood approach that allows us to estimate parameters for general models of network growth that can be expressed in terms of recursion relations. To handle larger networks we have developed an importance sampling scheme that allows us to approximate the likelihood and draw inference about the network and how it has been generated, estimate the parameters in the model, and perform parametric bootstrap analysis of network data. We illustrate the power of this approach by estimating growth parameters for the Caenorhabditis elegans protein interaction network.

Animals↗

Revising regulatory networks: from expression data to linear causal models.

Discovering the complex regulatory networks that govern mRNA expression is an important but difficult problem. Many current approaches use only expression data from microarrays to infer the likely network structure. However, this ignores much existing knowledge because for a given organism and system under study, a biologist may already have a partial model of gene regulation. We propose a method for revising and improving these initial models, which may be incomplete or partially incorrect, with expression data. We demonstrate our approach by revising a model of photosynthesis regulation proposed by a biologist for Cyanobacteria. Applied to wild type expression data, our system suggested several modifications consistent with biological knowledge. Applied to a mutant strain, our system correctly modified the disabled gene. Power experiments with synthetic data that indicate that reliable revision is feasible even with a small number of samples.

Gene Expression Profiling↗

Inferring topology from clustering coefficients in protein-protein interaction networks.

BACKGROUND: Although protein-protein interaction networks determined with high-throughput methods are incomplete, they are commonly used to infer the topology of the complete interactome. These partial networks often show a scale-free behavior with only a few proteins having many and the majority having only a few connections. Recently, the possibility was suggested that this scale-free nature may not actually reflect the topology of the complete interactome but could also be due to the error proneness and incompleteness of large-scale experiments. RESULTS: In this paper, we investigate the effect of limited sampling on average clustering coefficients and how this can help to more confidently exclude possible topology models for the complete interactome. Both analytical and simulation results for different network topologies indicate that partial sampling alone lowers the clustering coefficient of all networks tremendously. Furthermore, we extend the original sampling model by also including spurious interactions via a preferential attachment process. Simulations of this extended model show that the effect of wrong interactions on clustering coefficients depends strongly on the skewness of the original topology and on the degree of randomness of clustering coefficients in the corresponding networks. CONCLUSION: Our findings suggest that the complete interactome is either highly skewed such as e.g. in scale-free networks or is at least highly clustered. Although the correct topology of the interactome may not be inferred beyond any reasonable doubt from the interaction networks available, a number of topologies can nevertheless be excluded with high confidence.

Cluster Analysis↗

Applying dynamic Bayesian networks to perturbed gene expression data.

BACKGROUND: A central goal of molecular biology is to understand the regulatory mechanisms of gene transcription and protein synthesis. Because of their solid basis in statistics, allowing to deal with the stochastic aspects of gene expressions and noisy measurements in a natural way, Bayesian networks appear attractive in the field of inferring gene interactions structure from microarray experiments data. However, the basic formalism has some disadvantages, e.g. it is sometimes hard to distinguish between the origin and the target of an interaction. Two kinds of microarray experiments yield data particularly rich in information regarding the direction of interactions: time series and perturbation experiments. In order to correctly handle them, the basic formalism must be modified. For example, dynamic Bayesian networks (DBN) apply to time series microarray data. To our knowledge the DBN technique has not been applied in the context of perturbation experiments. RESULTS: We extend the framework of dynamic Bayesian networks in order to incorporate perturbations. Moreover, an exact algorithm for inferring an optimal network is proposed and a discretization method specialized for time series data from perturbation experiments is introduced. We apply our procedure to realistic simulations data. The results are compared with those obtained by standard DBN learning techniques. Moreover, the advantages of using exact learning algorithm instead of heuristic methods are analyzed. CONCLUSION: We show that the quality of inferred networks dramatically improves when using data from perturbation experiments. We also conclude that the exact algorithm should be used when it is possible, i.e. when considered set of genes is small enough.

Algorithms↗

A hybrid neural network algorithm for on-line state inference that accounts for differences in inoculum of Cephalosporium acremonium in fed-batch fermentors.

One serious difficulty in modeling a fermentative process is the forecasting of the duration of the lag phase. The usual approach to model biochemical reactors relies on first-principles, unstructured mathematical models. These models are not able to take into account changes in the process response caused by different incubation times or by repeated fedbatches. To overcome this problem, we have proposed a hybrid neural network algorithm. Feedforward neural networks were used to estimate rates of cell growth, substrate consumption, and product formation from on-line measurements during cephalosporin C production. These rates were included in the mass balance equations to estimate key process variables: concentrations of cells, substrate, and product. Data from fed-batch fermentation runs in a stirred aerated bioreactor employing the microorganism Cephalosporium acremonium ATCC 48272 were used. On-line measurements strongly related to the mass and activity of the cells used. They include carbon dioxide and oxygen concentrations in the exhausted gas. Good results were obtained using this approach.

Acremonium↗

transfactor: transcription factor activity estimation via probabilistic gene expression deconvolution.

Gene expression is a primary modality being studied to differentiate between biological cells. Contemporary single-cell studies simultaneously measure genome-wide transcription levels for thousands of individual cells in a single experiment. While the characterization of cell population differences has often occurred through differential gene expression analysis, tiny effect sizes become statistically significant when thousands of cells are available for each population, compromising biological interpretation. Moreover, these large studies have spurred the development of methods to infer gene regulatory networks (GRNs) directly from the data, and GRN databases are becoming more comprehensive. In this work, we propose a statistical model for gene expression measures and an inference method that leverage GRNs to deconvolve transcription factor (TF) activity from gene expression, by probabilistically assigning mRNA molecules to TFs. This shifts the paradigm from investigating gene expression differences to regulatory differences at the level of TF activity, aiding interpretation and allowing prioritization of a limited number of TFs responsible for significant contributions to the observed gene expression differences. The inferred TF activities result in intuitive prioritization of TFs in terms of the (difference in) estimated number of molecules they produce, in contrast to other widely used methods relying on arbitrary enrichment scores. Our model allows the incorporation of prior information on the regulatory potential between each TF and target gene and is able to deal with both repressing and activating interactions. We compare our approach to other TF activity estimation methods using two simulation experiments and two case studies. Single-cell RNA-sequencing; TF activity; bioinformatics; GRN.

Transcription Factors↗

Three machine learning techniques for automatic determination of rules to control locomotion.

Automatic prediction of gait events (e.g., heel contact, flat foot, initiation of the swing, etc.) and corresponding profiles of the activations of muscles is important for real-time control of locomotion. This paper presents three supervised machine learning (ML) techniques for prediction of the activation patterns of muscles and sensory data, based on the history of sensory data, for walking assisted by a functional electrical stimulation (FES). Those ML's are: 1) a multilayer perceptron with Levenberg-Marquardt modification of backpropagation learning algorithm; 2) an adaptive-network-based fuzzy inference system (ANFIS); and 3) a combination of an entropy minimization type of inductive learning (IL) technique and a radial basis function (RBF) type of artificial neural network with orthogonal least squares learning algorithm. Here we show the prediction of the activation of the knee flexor muscles and the knee joint angle for seven consecutive strides based on the history of the knee joint angle and the ground reaction forces. The data used for training and testing of ML's was obtained from a simulation of walking assisted with an FES system [39]. The ability of generating rules for an FES controller was selected as the most important criterion when comparing the ML's. Other criteria such as generalization of results, computational complexity, and learning rate were also considered. The minimal number of rules and the most explicit and comprehensible rules were obtained by ANFIS. The best generalization was obtained by the IL and RBF network.

Algorithms↗

A fuzzy logic-controlled classifier for use in implantable cardioverter defibrillators.

PURPOSE: Implantable cardioverters defibrillators (ICDs) are increasingly used in the management of life-threatening arrhythmias. Correct recognition of a treatable arrhythmia is crucial to this application. However, the computational power of microprocessors currently used in ICDs limits the range of traditional algorithms available for this application. METHODS: Classification based on fuzzy inference systems (FIS) were trained to recognize different cardiac rhythms (AF, VF, SVT, VT) from the Ann Arbor Electrogram Library. The FIS used were designed using adaptive-network-based fuzzy inference methods to optimize the classification procedure. Only computational techniques suitable for ICD design were used. RESULTS: After pretraining with the ANFIS correct rhythm classification was observed for the rhythms studied. CONCLUSION: In this preliminary study, successful rhythm classification was demonstrated using fuzzy logic techniques. In view of the computational efficiency this may have application in ICD design.

Arrhythmias, Cardiac↗

Computational strategy for discovering druggable gene networks from genome-wide RNA expression profiles.

We propose a computational strategy for discovering gene networks affected by a chemical compound. Two kinds of DNA microarray data are assumed to be used: One dataset is short time-course data that measure responses of genes following an experimental treatment. The other dataset is obtained by several hundred single gene knock-downs. These two datasets provide three kinds of information; (i) A gene network is estimated from time-course data by the dynamic Bayesian network model, (ii) Relationships between the knocked-down genes and their regulatees are estimated directly from knock-down microarrays and (iii) A gene network can be estimated by gene knock-down data alone using the Bayesian network model. We propose a method that combines these three kinds of information to provide an accurate gene network that most strongly relates to the mode-of-action of the chemical compound in cells. This information plays an essential role in pharmacogenomics. We illustrate this method with an actual example where human endothelial cell gene networks were generated from a novel time course of gene expression following treatment with the drug fenofibrate, and from 270 novel gene knock-downs. Finally, we succeeded in inferring the gene network related to PPAR-alpha, which is a known target of fenofibrate.

Bayes Theorem↗

Traveling waves of HIV infection on a low dimensional 'socio-geographic' network.

Observation of an essentially linear growth in time of U.S. and New York City AIDS cases, from about 1984 through early 1988, is shown to imply a relatively constant rate of transmission of HIV infection in its early stages, as has been observed for limited times in cohorts of male homosexuals in San Francisco and New York City. Observation by Potterat et al. of an exceptionally close intertwining of spatial and social patterns of endemic gonorrhea within a minority population, coupled with a percolation process model of HIV transmission within geographically constrained social networks, leads to inference that a constant rate of HIV transmission, in turn, implies a 'surface growth' phenomenon resulting in a traveling wave of infection advancing at a fixed 'velocity' along a 'one dimensional socio-geographic network.' Implications of this view are discussed for both data collection and analysis, and for intervention. Differences for the processes of disease transmission and control, based on the relative stability of socio-geographic networks, are postulated between the ghettoes of the middle-class male homosexual community and the physically devastated and socially distintegrated ghettoes of the minority urban poor.

Acquired Immunodeficiency Syndrome↗

Mathematical approaches to differentiation and gene regulation.

We consider some mathematical issues raised by the modelling of gene networks. The expression of genes is governed by a complex set of regulations, which is often described symbolically by interaction graphs. These are finite oriented graphs where vertices are the genes involved in the biological system of interest and arrows describe their interactions: a positive (resp. negative) arrow from a gene to another represents an activation (resp. inhibition) of the expression of the latter gene by some product of the former. Once such an interaction graph has been established, there remains the difficult task to decide which dynamical properties of the gene network can be inferred from it, in the absence of precise quantitative data about their regulation. There mathematical tools, among others, can be of some help. In this paper we discuss a rule proposed by Thomas according to which the possibility for the network to have several stationary states implies the existence of a positive circuit in the corresponding interaction graph. We prove that, when properly formulated in rigorous terms, this rule becomes a theorem valid for several different types of formal models of gene networks. This result is already known for models of differential [C. Soulé, Graphic requirements for multistationarity, ComPlexUs 1 (2003) 123-133] or Boolean [E. Rémy, P. Ruet, D. Thieffry, Graphic requirements for multistability and attractive cycles in a boolean dynamical framework, 2005, Preprint] type. We show here that a stronger version of it holds in the differential setup when the decay of protein concentrations is taken into account. This allows us to verify also the validity of Thomas' rule in the context of piecewise-linear models. We then discuss open problems.

Cell Differentiation↗

EXAMINE: a computational approach to reconstructing gene regulatory networks.

Reverse-engineering of gene networks using linear models often results in an underdetermined system because of excessive unknown parameters. In addition, the practical utility of linear models has remained unclear. We address these problems by developing an improved method, EXpression Array MINing Engine (EXAMINE), to infer gene regulatory networks from time-series gene expression data sets. EXAMINE takes advantage of sparse graph theory to overcome the excessive-parameter problem with an adaptive-connectivity model and fitting algorithm. EXAMINE also guarantees that the most parsimonious network structure will be found with its incremental adaptive fitting process. Compared to previous linear models, where a fully connected model is used, EXAMINE reduces the number of parameters by O(N), thereby increasing the chance of recovering the underlying regulatory network. The fitting algorithm increments the connectivity during the fitting process until a satisfactory fit is obtained. We performed a systematic study to explore the data mining ability of linear models. A guideline for using linear models is provided: If the system is small (3-20 elements), more than 90% of the regulation pathways can be determined correctly. For a large-scale system, either clustering is needed or it is necessary to integrate information in addition to expression profile. Coupled with the clustering method, we applied EXAMINE to rat central nervous system development (CNS) data with 112 genes. We were able to efficiently generate regulatory networks with statistically significant pathways that have been predicted previously.

Algorithms↗

A systematic approach to infer biological relevance and biases of gene network structures.

The development of high-throughput technologies has generated the need for bioinformatics approaches to assess the biological relevance of gene networks. Although several tools have been proposed for analysing the enrichment of functional categories in a set of genes, none of them is suitable for evaluating the biological relevance of the gene network. We propose a procedure and develop a web-based resource (BIOREL) to estimate the functional bias (biological relevance) of any given genetic network by integrating different sources of biological information. The weights of the edges in the network may be either binary or continuous. These essential features make our web tool unique among many similar services. BIOREL provides standardized estimations of the network biases extracted from independent data. By the analyses of real data we demonstrate that the potential application of BIOREL ranges from various benchmarking purposes to systematic analysis of the network biology.

Computational Biology↗

Variational learning for switching state-space models.

We introduce a new statistical model for time series that iteratively segments data into regimes with approximately linear dynamics and learnsthe parameters of each of these linear regimes. This model combines and generalizes two of the most widely used stochastic time-series models -- hidden Markov models and linear dynamical systems -- and is closely related to models that are widely used in the control and econometrics literatures. It can also be derived by extending the mixture of experts neural network (Jacobs, Jordan, Nowlan, & Hinton, 1991) to its fully dynamical version, in which both expert and gating networks are recurrent. Inferring the posterior probabilities of the hidden states of this model is computationally intractable, and therefore the exact expectation maximization (EM) algorithm cannot be applied. However, we present a variational approximation that maximizes a lower bound on the log-likelihood and makes use of both the forward and backward recursions for hidden Markov models and the Kalman filter recursions for linear dynamical systems. We tested the algorithm on artificial data sets and a natural data set of respiration force from a patient with sleep apnea. The results suggest that variational approximations are a viable method for inference and learning in switching state-space models.

Algorithms↗

Generation and propagation of subthreshold waves in a network of inferior olivary neurons.

The cells of the inferior olivary (IO) nucleus generate a large repertoire of electrical signals, among them subthreshold oscillations of the membrane potential (STO). To date, subthreshold oscillations have been studied at the level of single-cell recordings, from which network properties were inferred. In this study we used whole cell patch recordings and optical imaging to address the following issues: 1) synchrony of STO in neighboring neurons; 2) stability of the oscillatory activity in the temporal and spatial domain; and 3) the size of the oscillating network. Recordings were made from 126 pairs of IO neurons in 13- to 30-day-old rats. An additional 262 neurons were recorded individually. The frequency of STO varied from 0.8 to 8.6 Hz. The frequency distribution revealed two subpopulations with peaks at about 3 and 6 Hz. The maximum amplitude among the cells varied from 2 to 25 mV. Oscillations in most neurons showed ongoing modulations in both frequency and amplitude. These modulations were largely abolished following bath application of 40 microM 6-cyano-7-nitroquinoxaline-2,3-dione (CNQX), a competitive non-N-methyl-D-aspartate (non-NMDA) receptor antagonist, suggesting that they were caused by glutamatergic action. In 35 of 61 recorded pairs at least one neuron exhibited STO permitting us to compare frequency and phase relations. In 22 pairs there was coherent activity with zero phase difference between oscillations in the 2 cells. In these pairs, frequency and amplitude modulation occurred simultaneously in both neurons. Electrotonic coupling was tested in 13 pairs, that had coherent STO, and it was detected in 12. An additional seven pairs showed coherent oscillations but with a phase difference of 20-50 ms. Electrotonic coupling was observed in three of these pairs. Electrotonic coupling was also observed in two of five pairs in which only one neuron oscillated. No coupling was detected in one pair where both neurons oscillated but at different frequencies. Optical imaging using a voltage-sensitive dye (RH 414) was performed on 40 IO slices using an array of 128 photodiodes. Patches of oscillatory activity were observed in 10 slices. Among them six showed spontaneous oscillations, and four exhibited oscillations following extracellular stimulation. In agreement with cell pair recording, optical imaging demonstrated phase-shifted activity in the form of propagating waves of activity within an oscillating patch. We conclude that 1) STO exhibit ongoing modulations of frequency and amplitude that are probably caused by extrinsic inputs to the IO nucleus; 2) electrotonically coupled neurons show a high level of STO synchrony; and 3) the oscillatory activity can propagate within a network of coupled olivary neurons.

6-Cyano-7-nitroquinoxaline-2,3-dione↗

Sonohistology for the computerized differentiation of parotid gland tumors.

A system for the computerized differentiation of parotid gland tumors is proposed. The parotid gland is the largest of the salivary glands. It is found in the subcutaneous tissue of the face, overlying the mandibular ramus and anterior and inferior to the external ear. The classification system is based on a multifeature tissue characterization approach, using fuzzy inference systems as higher-order classifiers. Baseband ultrasonic echo data were acquired during conventional ultrasound imaging examinations using standard ultrasound equipment. Several tissue-describing parameters were calculated within numerous small regions of interest to evaluate spectral and textural tissue properties. The parameters were processed by an adaptive network-based fuzzy inference system, using the results of conventional histology after parotidectomy as the "gold standard." The results of the classification are presented as a numerical score indicating the probability of a certain tumor or alteration for each parotid gland. The score can be presented to the physician during examination of the patient to improve the differentiation between various types of parotid gland tumors. The system was evaluated on n = 23 cases of patients undergoing radical parotidectomy. The receiver operating characteristic curve area is A(ROC) = 0.95 +/- 0.07 when using fourfold cross-validation over cases and differentiating between various benign parotid gland tumors and monomorphic adenoma.

Adenoma↗