PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Network inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Genome-wide SNP data support species boundaries in sympatric Polylepis Ruiz & Pav. (Rosaceae) species from Bolivia and Ecuador.

Species delimitation in the South American genus Polylepis is notoriously challenging due to high morphological similarity and phenotypic plasticity, likely driven by hybridization and gene flow. Previous phylogenetic studies suggested that genetic structure aligns more strongly with geography than with taxonomy, questioning existing species concepts and hampering conservation efforts. We used double-digest RAD sequencing (ddRADseq) to generate genome-wide SNP data for 11 Polylepis species sampled across multiple localities in Bolivia and Ecuador. Population genetic analyses, phylogenetic inference, and network approaches were combined to assess whether genetic structure aligns more closely with taxonomy or geography. Morphologically defined species formed largely cohesive genetic lineages across regions, with species identity explaining substantially more genetic variation than locality. While localized admixture and reticulation were detected among closely related taxa, widespread species showed strong genetic cohesion and clear separation from congeners. Our results indicate that the sampled Polylepis species from Bolivia and Ecuador maintain distinct genetic identities despite localized signals consistent with gene flow. This genome-wide support for current taxonomy highlights Polylepis as a valuable model for studying speciation under gene flow and indicates that multiple geographic sampling will be essential in reconstructing a robust phylogeny of the genus, with important implications for conservation planning in Andean montane forests.

Bolivia↗

Fuzzy logic model of Langmuir probe discharge data.

Plasma models are crucial to gain physical insights into complex discharges as well as to optimizing plasma-driven processes. As an alternative to physical model, a qualitative model was constructed using adaptive fuzzy logic called adaptive network fuzzy inference system (ANFIS). Prediction performance of ANFIS was evaluated on two sets of experimental discharge data. One referred to as hemispherical inductively coupled plasma (HICP) was characterized with a 2(4) full factorial experiment, in which the factors that were varied include source power, pressure, chuck position, and Cl2 flow rate. The other called multipole ICP was characterized by performing a 3(3) full factorial experiment on the factors, including source power, pressure, and Ar flow rate. Trained ANFIS models were tested on eight and 16 experiments not pertaining to previous training data for HICP and MICP, respectively. Plasma attributes modeled include electron density. electron temperature, and plasma potential. The performance of ANFIS was optimized as a function of a type of membership function, number of membership function, and two learning factors. The number of membership functions was different depending on the type of plasma data and employing too large number of membership functions resulted in a drastic degradation in prediction performances. Optimized ANFIS models were compared to statistical regression models and demonstrated improved predictions in all comparisons.

Journal Article↗

A single-nucleus transcriptomic atlas of human inner ear development.

Hearing and balance rely on coordinated activity of multiple inner ear cell types, yet the mechanisms governing their development and specification in humans remain unclear. Consequently, this limits our understanding of how disease genes affect cell type formation and function, limiting the development of targeted treatments, including gene therapies. Here we present the Human Inner Ear Development snRNA-seq Atlas (HIEDRA), a single-nucleus transcriptomic atlas of the human inner ear spanning the first and second trimesters. HIEDRA maps sensory and nonsensory epithelia, neurons and mesenchyme-associated populations, including undercharacterized secretory cells required for ion homeostasis. We identify selective vulnerability in sensory and secretory lineages to disease-associated genes, infer regulatory networks and show that Hedgehog signaling suppression is required for secretory cell specification. We validate this mechanism in human inner ear organoids, expanding the model to include all major cell types. Altogether, these findings provide insights into human inner ear cell type specification, improve in vitro models and establish HIEDRA as a resource for investigating human inner ear development.

Journal Article↗

Extended linkage disequilibrium surrounding the hemoglobin E variant due to malarial selection.

The hemoglobin E variant (HbE; ( beta )26Glu-->Lys) is concentrated in parts of Southeast Asia where malaria is endemic, and HbE carrier status has been shown to confer some protection against Plasmodium falciparum malaria. To examine the effect of natural selection on the pattern of linkage disequilibrium (LD) and to infer the evolutionary history of the HbE variant, we analyzed biallelic markers surrounding the HbE variant in a Thai population. Pairwise LD analysis of HbE and 43 surrounding biallelic markers revealed LD of HbE extending beyond 100 kb, whereas no LD was observed between non-HbE variants and the same markers. The inferred haplotype network suggests a single origin of the HbE variant in the Thai population. Forward-in-time computer simulations under a variety of selection models indicate that the HbE variant arose 1,240-4,440 years ago. These results support the conjecture that the HbE mutation occurred recently, and the allele frequency has increased rapidly. Our study provides another clear demonstration that a high-resolution LD map across the human genome can detect recent variants that have been subjected to positive selection.

Animals↗

VirBinn improves viral genome binning from metagenomic Hi-C through graph diffusion.

MOTIVATION: Metagenomic Hi-C provides in situ proximity signals that can improve genome binning and enable virus-host-association analysis. However, viral genome recovery remains difficult because virus-virus Hi-C contact matrices are extremely sparse. Viral genomes are small, often low-abundance, and frequently assemble into short contigs, leaving many true within-genome links unobserved and causing viral bins to fragment. RESULTS: We present VirBinn, a graph-diffusion framework for viral binning from metagenomic Hi-C. VirBinn enhances virus-virus connectivity through two complementary mechanisms: random-walk-with-restart enhancement on the sparse virus-virus contact graph and host-guided diffusion that propagates viral seeds through the host network to infer indirect virus-virus associations. The enhanced views are integrated and clustered using Leiden community detection to produce viral metagenome-assembled genomes (vMAGs). On dataset-specific simulation benchmarks with ground truth, VirBinn consistently recovers more high-quality vMAGs than Hi-C-based and shotgun-based baselines and substantially increases the number of near-complete genomes. On four real metagenomic Hi-C datasets spanning human gut, pig gut, sheep gut (long-read assembly), and wastewater, VirBinn yields more high-completeness vMAGs under CheckV and produces bins with strong within-cluster contact support. Finally, host linkage analysis using reconstructed host MAGs reveals habitat-specific host-association patterns and plausible host taxonomic profiles. AVAILABILITY AND IMPLEMENTATION: VirBinn is available at https://github.com/dyxstat/VirBinn. The scripts to reproduce the results and figures in this article are available at https://github.com/dyxstat/Reproduce_VirBinn.

Genome, Viral↗

Inferring quantitative models of regulatory networks from expression data.

MOTIVATION: Genetic networks regulate key processes in living cells. Various methods have been suggested to reconstruct network architecture from gene expression data. However, most approaches are based on qualitative models that provide only rough approximations of the underlying events, and lack the quantitative aspects that are critical for understanding the proper function of biomolecular systems. RESULTS: We present fine-grained dynamical models of gene transcription and develop methods for reconstructing them from gene expression data within the framework of a generative probabilistic model. Unlike previous works, we employ quantitative transcription rates, and simultaneously estimate both the kinetic parameters that govern these rates, and the activity levels of unobserved regulators that control them. We apply our approach to expression datasets from yeast and show that we can learn the unknown regulator activity profiles, as well as the binding affinity parameters. We also introduce a novel structure learning algorithm, and demonstrate its power to accurately reconstruct the regulatory network from those datasets.

Binding Sites↗

An evolutionary approach for gene expression patterns.

This study presents an evolutionary algorithm, called a heterogeneous selection genetic algorithm (HeSGA), for analyzing the patterns of gene expression on microarray data. Microarray technologies have provided the means to monitor the expression levels of a large number of genes simultaneously. Gene clustering and gene ordering are important in analyzing a large body of microarray expression data. The proposed method simultaneously solves gene clustering and gene-ordering problems by integrating global and local search mechanisms. Clustering and ordering information is used to identify functionally related genes and to infer genetic networks from immense microarray expression data. HeSGA was tested on eight test microarray datasets, ranging in size from 147 to 6221 genes. The experimental clustering and visual results indicate that HeSGA not only ordered genes smoothly but also grouped genes with similar gene expressions. Visualized results and a new scoring function that references predefined functional categories were employed to confirm the biological interpretations of results yielded using HeSGA and other methods. These results indicate that HeSGA has potential in analyzing gene expression patterns.

Algorithms↗

In vivo contribution of h-channels in the septal pacemaker to theta rhythm generation.

One of the most intriguing network-level inferences made on the basis of in vitro and modelling data regarding the role of Ih current was that they participate in rhythmogenesis in different parts of the brain. The nature of Ih contribution to various neuronal oscillations is far from uniform however, and the proper evaluation of the role of Ih in each particular structure requires in situ investigations in the intact brain. In this study we tested the effect of Ih blockade in the medial septum on hippocampal theta rhythm in anaesthetized and freely behaving rats. We could not confirm the recent report of elimination of theta by septal injection of ZD7288 [C. Xu et al. (2004) Eur. J. Neurosci., 19, 2299-2309]; the observed effects were more subtle and more specific. We found that Ih blockade in the medial septum substantially decreased the frequency of hippocampal oscillations without changing the context in which theta occurred, i.e. specific behaviours in freely moving rats and spontaneous switching and brainstem stimulation under anaesthesia. Septal injection of ZD7288 eliminated atropine-resistant theta elicited by high intensity electrical stimulation of the reticular formation in anaesthetized rats but was ineffective in combination with the muscarinic agonist, carbachol. Thus, functional Ih was necessary for the septum to generate or transmit high frequency theta rhythm elicited by strong ascending activation, whereas low frequency theta persisted after Ih blockade. These results suggest that Ih plays a specific role in septal theta generation by promoting fast oscillations during exploratory behaviour and rapid eye movement sleep.

Action Potentials↗

Unique transcriptome signature of Mycobacterium tuberculosis in pulmonary tuberculosis.

Although tuberculosis remains a substantial global threat, the mechanisms that enable mycobacterial persistence and replication within the human host are ill defined. This study represents the first genome-wide expression analysis of Mycobacterium tuberculosis from clinical lung samples, which has enabled the identification of M. tuberculosis genes actively expressed during pulmonary tuberculosis. To obtain optimal information from our DNA array analyses, we analyzed the differentially expressed genes within the context of computationally inferred protein networks. Protein networks were constructed using functional linkages established by the Rosetta stone, phylogenetic profile, conserved gene neighbor, and operon computational methods. This combined approach revealed that during pulmonary tuberculosis, M. tuberculosis actively transcribes a number of genes involved in active fortification and evasion from host defense systems. These genes may provide targets for novel intervention strategies.

Bacterial Proteins↗

Sequential RAM-based neural networks: learnability, generalisation, knowledge extraction, and grammatical inference.

A fundamental question in the field of artificial neural networks is what set of problems a given class of networks can perform (computability). Such a problem can be made less general, but no less important, by asking what these networks could learn by using a given training procedure (learnability). The basic purpose of this paper is to address the learnability problem. Specifically, it analyses the learnability of sequential RAM-based neural networks. The analytical tools used are those of Automata Theory. In this context, this paper establishes which class of problems and under what conditions such networks, together with their existing learning rules, can learn and generalize. This analysis also yields techniques for both extracting knowledge from and inserting knowledge into the networks. The results presented here, besides helping in a better understanding of the temporal behaviour of sequential RAM-based networks, could also provide useful insights for the integration of the symbolic/connectionist paradigms.

Algorithms↗

Analysis of molecular profile data using generative and discriminative methods.

A modular framework is proposed for modeling and understanding the relationships between molecular profile data and other domain knowledge using a combination of generative (here, graphical models) and discriminative [Support Vector Machines (SVMs)] methods. As illustration, naive Bayes models, simple graphical models, and SVMs were applied to published transcription profile data for 1,988 genes in 62 colon adenocarcinoma tissue specimens labeled as tumor or nontumor. These unsupervised and supervised learning methods identified three classes or subtypes of specimens, assigned tumor or nontumor labels to new specimens and detected six potentially mislabeled specimens. The probability parameters of the three classes were utilized to develop a novel gene relevance, ranking, and selection method. SVMs trained to discriminate nontumor from tumor specimens using only the 50-200 top-ranked genes had the same or better generalization performance than the full repertoire of 1,988 genes. Approximately 90 marker genes were pinpointed for use in understanding the basic biology of colon adenocarcinoma, defining targets for therapeutic intervention and developing diagnostic tools. These potential markers highlight the importance of tissue biology in the etiology of cancer. Comparative analysis of molecular profile data is proposed as a mechanism for predicting the physiological function of genes in instances when comparative sequence analysis proves uninformative, such as with human and yeast translationally controlled tumour protein. Graphical models and SVMs hold promise as the foundations for developing decision support systems for diagnosis, prognosis, and monitoring as well as inferring biological networks.

Bayes Theorem↗

Finding groups in gene expression data.

The vast potential of the genomic insight offered by microarray technologies has led to their widespread use since they were introduced a decade ago. Application areas include gene function discovery, disease diagnosis, and inferring regulatory networks. Microarray experiments enable large-scale, high-throughput investigations of gene activity and have thus provided the data analyst with a distinctive, high-dimensional field of study. Many questions in this field relate to finding subgroups of data profiles which are very similar. A popular type of exploratory tool for finding subgroups is cluster analysis, and many different flavors of algorithms have been used and indeed tailored for microarray data. Cluster analysis, however, implies a partitioning of the entire data set, and this does not always match the objective. Sometimes pattern discovery or bump hunting tools are more appropriate. This paper reviews these various tools for finding interesting subgroups.

Journal Article↗

Design of microarray experiments for genetical genomics studies.

Microarray experiments have been used recently in genetical genomics studies, as an additional tool to understand the genetic mechanisms governing variation in complex traits, such as for estimating heritabilities of mRNA transcript abundances, for mapping expression quantitative trait loci, and for inferring regulatory networks controlling gene expression. Several articles on the design of microarray experiments discuss situations in which treatment effects are assumed fixed and without any structure. In the case of two-color microarray platforms, several authors have studied reference and circular designs. Here, we discuss the optimal design of microarray experiments whose goals refer to specific genetic questions. Some examples are used to illustrate the choice of a design for comparing fixed, structured treatments, such as genotypic groups. Experiments targeting single genes or chromosomic regions (such as with transgene research) or multiple epistatic loci (such as within a selective phenotyping context) are discussed. In addition, microarray experiments in which treatments refer to families or to subjects (within family structures or complex pedigrees) are presented. In these cases treatments are more appropriately considered to be random effects, with specific covariance structures, in which the genetic goals relate to the estimation of genetic variances and the heritability of transcriptional abundances.

Animals↗

Calculating the statistical significance of changes in pathway activity from gene expression data.

We present a statistical approach to scoring changes in activity of metabolic pathways from gene expression data. The method identifies the biologically relevant pathways with corresponding statistical significance. Based on gene expression data alone, only local structures of genetic networks can be recovered. Instead of inferring such a network, we propose a hypothesis-based approach. We use given knowledge about biological networks to improve sensitivity and interpretability of findings from microarray experiments. Recently introduced methods test if members of predefined gene sets are enriched in a list of top-ranked genes in a microarray study. We improve this approach by defining scores that depend on all members of the gene set and that also take pairwise co-regulation of these genes into account. We calculate the significance of co-regulation of gene sets with a nonparametric permutation test. On two data sets the method is validated and its biological relevance is discussed. It turns out that useful measures for co-regulation of genes in a pathway can be identified adaptively. We refine our method in two aspects specific to pathways. First, to overcome the ambiguity of enzyme-to-gene mappings for a fixed pathway, we introduce algorithms for selecting the best fitting gene for a specific enzyme in a specific condition. In selected cases, functional assignment of genes to pathways is feasible. Second, the sensitivity of detecting relevant pathways is improved by integrating information about pathway topology. The distance of two enzymes is measured by the number of reactions needed to connect them, and enzyme pairs with a smaller distance receive a higher weight in the score calculation.

Journal Article↗

Construction of phylogenetic trees by kernel-based comparative analysis of metabolic networks.

BACKGROUND: To infer the tree of life requires knowledge of the common characteristics of each species descended from a common ancestor as the measuring criteria and a method to calculate the distance between the resulting values of each measure. Conventional phylogenetic analysis based on genomic sequences provides information about the genetic relationships between different organisms. In contrast, comparative analysis of metabolic pathways in different organisms can yield insights into their functional relationships under different physiological conditions. However, evaluating the similarities or differences between metabolic networks is a computationally challenging problem, and systematic methods of doing this are desirable. Here we introduce a graph-kernel method for computing the similarity between metabolic networks in polynomial time, and use it to profile metabolic pathways and to construct phylogenetic trees. RESULTS: To compare the structures of metabolic networks in organisms, we adopted the exponential graph kernel, which is a kernel-based approach with a labeled graph that includes a label matrix and an adjacency matrix. To construct the phylogenetic trees, we used an unweighted pair-group method with arithmetic mean, i.e., a hierarchical clustering algorithm. We applied the kernel-based network profiling method in a comparative analysis of nine carbohydrate metabolic networks from 81 biological species encompassing Archaea, Eukaryota, and Eubacteria. The resulting phylogenetic hierarchies generally support the tripartite scheme of three domains rather than the two domains of prokaryotes and eukaryotes. CONCLUSION: By combining the kernel machines with metabolic information, the method infers the context of biosphere development that covers physiological events required for adaptation by genetic reconstruction. The results show that one may obtain a global view of the tree of life by comparing the metabolic pathway structures using meta-level information rather than sequence information. This method may yield further information about biological evolution, such as the history of horizontal transfer of each gene, by studying the detailed structure of the phylogenetic tree constructed by the kernel-based method.

Archaeal Proteins↗

The effects of incomplete protein interaction data on structural and evolutionary inferences.

BACKGROUND: Present protein interaction network data sets include only interactions among subsets of the proteins in an organism. Previously this has been ignored, but in principle any global network analysis that only looks at partial data may be biased. Here we demonstrate the need to consider network sampling properties explicitly and from the outset in any analysis. RESULTS: Here we study how properties of the yeast protein interaction network are affected by random and non-random sampling schemes using a range of different network statistics. Effects are shown to be independent of the inherent noise in protein interaction data. The effects of the incomplete nature of network data become very noticeable, especially for so-called network motifs. We also consider the effect of incomplete network data on functional and evolutionary inferences. CONCLUSION: Crucially, when only small, partial network data sets are considered, bias is virtually inevitable. Given the scope of effects considered here, previous analyses may have to be carefully reassessed: ignoring the fact that present network data are incomplete will severely affect our ability to understand biological systems.

Evolution, Molecular↗

Comparison of recent methods for inference of variable influence in neural networks.

Neural networks (NNs) belong to 'black box' models and therefore 'suffer' from interpretation difficulties. Four recent methods inferring variable influence in NNs are compared in this paper. The methods assist the interpretation task during different phases of the modeling procedure. They belong to information theory (ITSS), the Bayesian framework (ARD), the analysis of the network's weights (GIM), and the sequential omission of the variables (SZW). The comparison is based upon artificial and real data sets of differing size, complexity and noise level. The influence of the neural network's size has also been considered. The results provide useful information about the agreement between the methods under different conditions. Generally, SZW and GIM differ from ARD regarding the variable influence, although applied to NNs with similar modeling accuracy, even when larger data sets sizes are used. ITSS produces similar results to SZW and GIM, although suffering more from the 'curse of dimensionality'.

Algorithms↗