PubMed HealthSearch

SEARCH · PubMed Health

Results for “Graph”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

HallmarkGraph: a cancer hallmark informed graph neural network for classifying hierarchical tumor subtypes.

MOTIVATION: Accurate tumor subtype diagnosis is crucial for precision oncology, yet current methodologies face significant challenges. These include balancing model accuracy with interpretability and the high costs of generating multi-omics data in clinical settings. Moreover, there is a lack of validated models capable of classifying hierarchical tumor subtypes across a comprehensive pan-cancer cohort. RESULTS: We present a graph neural network, HallmarkGraph, the first biologically informed model developed to classify hierarchical tumor subtypes in human cancer. Inspired by cancer hallmarks, the model's architecture integrates transcriptome profiles and gene regulatory interactions to perform multi-label classification. We evaluate the model on a comprehensive pan-cancer cohort comprising 11 476 samples from 26 primary cancers with 405 subtypes up to eight levels. The model demonstrates exceptional performance, achieving 5-fold cross-validation accuracy between 85% and 99% for tumor subtypes labeled with increasing details of genomic information. It also shows good generalizability on a validation dataset of 887 samples, assessed using three metrics that consider tumor subtypes at individual, combined, and sample levels. Benchmarking and ablation experiments show that hallmark-based embeddings slightly influence model performance, while the integrated multilayer perceptron plays a significant role in determining classifier accuracy. Additionally, we use the SHAP method to link cancer hallmarks with genes, identifying key features that influence model decisions. Our findings present a biologically informed machine learning framework capable of tracking tumor transcriptomic trajectories and distinguishing inter- and intra-tumor heterogeneity in pan-cancer. This approach holds promise for enhancing cancer diagnostics. AVAILABILITY AND IMPLEMENTATION: HallmarkGraph is accessible at https://github.com/laixn/HallmarkGraph.

Humans

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

A General Framework for Branch Length Estimation in Ancestral Recombination Graphs.

Inference of Ancestral Recombination Graphs (ARGs) is of central interest in the analysis of genomic variation. ARGs can be specified in terms of topologies and coalescence times. The coalescence times are usually estimated using an informative prior derived from coalescent theory, but this may generate biased estimates and can also complicate downstream inferences based on ARGs. Here we introduce, POLEGON, a novel approach for estimating branch lengths for ARGs which uses an uninformative prior. Using extensive simulations, we show that this method provides improved estimates of coalescence times and lead to more accurate inferences of effective population sizes under a wide range of demographic assumptions (population expansion, bottleneck, split, etc). It also improves other downstream inferences including estimates of mutation rates. We apply the method to data from the 1000 Genomes Project to investigate population size histories and differential mutation signatures across populations. We also estimate coalescence times in the HLA region, and show that they exceed 30 million years in multiple segments.

Ancestral Recombination Graph

SPC: a SPectral Component approach leveraging Identity-by-Descent graphs to address recent population structure in genomic analysis.

Population structure is a well-known confounder in statistical genetics, particularly in genome-wide association studies (GWAS), where it can lead to inflated test statistics and spurious associations. Traditional methods, such as principal components (PCs), commonly used to adjust for population structure, are limited in capturing fine-scale, non-linear patterns that arise from recent demographic events - patterns that are crucial for understanding rare variant effects. To address this challenge, we propose a novel method called SPectral Components (SPCs), which leverages identity-by-descent (IBD) graphs to capture and transform local, non-linear fine-scale population structure into continuous representations that can be seamlessly integrated into genetic analysis pipelines. Using both simulated datasets and empirical data from the UK Biobank (N ≈ 420,000), we demonstrate that SPCs outperform PCs in adjusting for fine-scale population structure. In simulations, SPCs explained over 90% of the fine-scale population structure with fewer components, while PCs captured less than 5%. In the UK Biobank, SPCs reduced the inflation of p-values in the GWAS of an environmental-driven phenotype by 12% compared to PCs, while maintaining a similar performance to PCs in height, a highly heritable phenotype. Additionally, SPCs improved rare variant association analyses, reducing genomic inflation (e.g., from 7.6 to 1.2 in one analysis), and provided more accurate heritability estimates. Spatial autocorrelation analysis further confirmed the ability of SPCs to account for environmental effects, reducing Moran's I for both environmental and heritable phenotypes more effectively than PCs. Overall, our findings demonstrate that SPCs provide a robust, scalable adjustment for recent population structure, offering a powerful alternative or complement to PCs in large-scale biobank studies.

GWAS

A graph-based approach for the visualisation and analysis of bacterial pangenomes.

BACKGROUND: The advent of low cost, high throughput DNA sequencing has led to the availability of thousands of complete genome sequences for a wide variety of bacterial species. Examining and interpreting genetic variation on this scale represents a significant challenge to existing methods of data analysis and visualisation. RESULTS: Starting with the output of standard pangenome analysis tools, we describe the generation and analysis of interactive, 3D network graphs to explore the structure of bacterial populations, the distribution of genes across a population, and the syntenic order in which those genes occur, in the new open-source network analysis platform, Graphia. Both the analysis and the visualisation are scalable to datasets of thousands of genome sequences. CONCLUSIONS: We anticipate that the approaches presented here will be of great utility to the microbial research community, allowing faster, more intuitive, and flexible interaction with pangenome datasets, thereby enhancing interpretation of these complex data.

Bacteria

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans

A graph analysis of the relationship between population density and social pathology.

This article deals with systematic variables--population density and pathological effects--at the level of a community or organizational system. Existing studies of the relationship between population density and social pathology (e.g., mortality, mental illness) in humans have failed to determine a strictly causal relationship between density and pathology in a linear model. They have not established a direct link between density and pathology in a simple two-variable, linear causal approach. Eliminating other variables by statistical control has obscured the conditions of a complex phenomenon. By incorporating socioeconomic variables into a general system model, this paper suggests that the total configuration of social organization, adaptation, previous group experience, and environment determine the effects of population density. Using a method of graph analysis, a model is presented of the relationship between population density and social pathology which has a high degree of isomorphism with the empirical situation.

Adrenocorticotropic Hormone

Bond graphs and the exploitation of power conserving transformations.

Use is made of the guaranteed energy conservation property of any bond graph model (presuming that consistent energy variables are used). Thus power transformations are power conserving, and this property is exploited with respect to multiport transformers where many effort and flow variables may be involved. In these cases the relationship among the effort (flow) variables across the transformer may be easier to derive than the relationship among the flows (efforts). The power conserving nature of the transformer permits immediate derivation of the alternate variable relationship. This formulation procedure is applied to the reflex reaction of the arm.

Energy Metabolism

Three-dimensional structural resemblance between leucine aminopeptidase and carboxypeptidase A revealed by graph-theoretical techniques.

Using 3-D searching techniques based on algorithms derived from graph theory we have established a striking structural similarity between the structure of bovine carboxypeptidase A and that of the C-terminal domain of bovine leucine aminopeptidase. There is no significant sequence homology between the aminopeptidases and the carboxypeptidases but the strong structural relationship detected in this complex fold suggests that there may be a very remote divergent evolutionary relationship between these two enzyme classes.

Animals

Amino acid transport across the human blood-CSF barrier. An evaluation graph for amino acid concentrations in cerebrospinal fluid.

The correlation of cerebrospinal fluid (CSF)/serum concentration quotients was used as a method for identification of amino acids which are transported by a common carrier system across the blood-CSF barrier. Isoleucine, leucine, valine, phenylalanine, tyrosine and lysine were found to compete for the same carrier system. This group of amino acids in man was found to be different from the system described as a neutral amino acid carrier at the blood-brain barrier in rats. In man, methionine and tryptophan do not compete with the other neutral amino acids for the same carrier system. In contrast, lysine as a basic amino acid is found to be correlated with the same transport system as the five neutral amino acids. A graph for the evaluation of pathological amino acid concentrations in CSF is presented. Patients with a blood-CSF barrier dysfunction for proteins showed partly normal, partly increased, CSF/serum concentration quotients for the amino acids. Hydroxyproline could be identified as a constituent of the amino acid pool in CSF. For proline and hydroxyproline a special control system has to be suggested because of their smaller biological variance in CSF than in blood. Contrary to the other amino acids proline and hydroxyproline have a smaller biological variation in CSF than in serum.

Adolescent

The corticosteroid dose graph. Use in determination of corticosteroid requirements, characterization of patients, and prevention of seasonal exacerbation in corticosteroid-dependent patients.

Sixty-three corticosteroid(CS)-dependent asthmatics were evaluated by a method which records asthmatic exacerbations as a function of change in CS dosage. This method, the Corticosteroid Dose Graph (CSDG), is a graphic representation of the patients' narrative charts. It is useful in assessing and directing the care of these patients and in evaluating the efficacy of altered therapeutic regimens in asthmatics. The CSDG categorizes cooperative patients (CP) and distinguishes these patients from those who are less cooperative (LCP). Evaluation of the progress of CP by the graphic analysis system showed that CP have fewer severe exacerbations of asthma than LCPs. Evaluation of the 63 patients by this graphic method shows a significantly increased incidence of asthmatic exacerbations in the Chicago area in October as compared to other months. The CSDG also demonstrates that some individuals experienced significantly increased numbers of asthma exacerbations at certain times of the year unique to themselves. Because the CSDG predicts exacerbations, it was shown to be applicable to prevent such exacerbations by prophylactic increase in CS dose during a high-risk season for selected CS-dependent asthmatics.

Adolescent

MEPS parameters and graph analysis for the use of recombination to construct ordered sets of overlapping clones.

Homologous recombination can provide a basis for the construction of an ordered set of overlapping clones. The principle is to make two libraries, each in a vector that has a different selectable marker flanking the insert site. Recombination between the flanking markers, leading to a selectable phenotype, can only occur as the consequence of crossing over between inserts. The two libraries are crossed in a matrix, allowing the construction of an ordered set. The logic, akin to S. Benzer's (1961, Genetics 47:403-415) for the arrangement of deletion and point mutations, has a graph theoretic formulation, which helps to cope with the complex and noisy data inherent in the physical mapping of genomes rich in repeated sequences. The minimum length of identity required for homologous recombination is called the MEPS (minimum efficient processing segment) and is a property of each recombination pathway. The amount and the type of sequence similarity required for two sequences to recombine is different from that implied by either the conservation of restriction sites or by most procedures of hybridization.

Chromosome Deletion

Robust error-minimization in the genetic code across physicochemical metrics and variant codes: A graph-theoretic analysis in GF(2)6.

The standard genetic code reduces the impact of point mutations, but the robustness of this property across physicochemical metrics, naturally occurring variant codes, and codon-reassignment mechanisms remains incompletely quantified. Embedding the 64 codons in GF(2)6 represents the hypercube Q6 as a coordinate-dependent subgraph of the encoding-independent single-nucleotide mutation graph H(3,4), and enables continuous &#x3c1;-interpolation between the two. Under a quartet-pattern shuffle null (n=10,000), the standard code is significantly low-cost across four established, code-independent physicochemical distance metrics with partially overlapping content (Grant ham p=0.0062; Miyata p<0.001; Woese polar requirement p=0.003; Kyte-Doolittle hydropathy p=0.001), and the signal strengthens monotonically as &#x3c1; moves Q6&#x2192;H(3,4). A structure-aware sensitivity analysis under the alignment-derived ProtSub matrix (Jia & Jernigan 2021) yields the most extreme percentile of any measure tested (p=0.0004; all five p-values pass Bonferroni at &#x3b1;=0.05). Across the 27 NCBI translation tables, near-optimality is preserved: 11 of 12 informative-distance variants retain top-5% placement after BH-FDR correction. Natural codon reassignments avoid disrupting codon-family connectivity: under the encoding-independent H(3,4) adjacency, observed events are topology-breaking at relative risk 0.32 versus the candidate landscape (permutation p&#x2264;10-4). The H(3,4) result is stable by construction; the Q6 decomposition is representation-specific and fails to show depletion under 8 of 24 base-to-bit encodings, so we report H(3,4) as the primary test and Q6 as a sensitivity. Event-level conditional-logit modelling shows that topology avoidance and local physicochemical cost provide complementary, only weakly correlated signal (rs=0.15), and that topology adds explanatory value beyond physicochemistry under both Q6 and encoding-independent H(3,4) adjacency. Retrospective reanalysis of nine genome-recoding datasets is consistent with codon-family topology operating as an evolutionary-trajectory constraint distinct from acute engineering fitness. The contribution is the second axis: code evolution is jointly constrained by physicochemical smoothness and codon-family topological integrity, and these two constraints are partly independent.

Codon reassignment

Propagation of information in MetaNet graph models.

Information flow in metabolic networks has been studied with a graph model which represents the biochemical transformations occurring in the system under investigation. The "signal strength", an algebraic expression which estimates the probability that an intermediate metabolite is bound to a given enzyme, has been used to derive the "signal transmittance", the fraction of the informational signal at one intermediate that reaches another intermediate. The transmittance has been used to derive the "response ratio", the sensitivity of the rate of change of information at one metabolite consequent to a perturbation at another metabolite. Because the graphical representation corresponds to the biochemical events presumed to occur in the network, these quantities can be used to design experiments to confirm or falsify the hypotheses underlying the model and aid in understanding the regulatory properties of the system. The technique is illustrated by an example model, and its predictions are shown to be sensitive to modest structural changes in the network.

Animals

Quantitative analysis of metabolic regulation. A graph-theoretic approach using spanning trees.

A graph-theoretic technique using spanning trees is described for the evaluation of Flux Control Coefficients of metabolic pathways. The technique is illustrated by investigating a linear pathway (a) in the absence of feedback and feedforward regulation. (b) with its first enzyme inhibited by the end product and (c) with multiple feedback loops. It is shown that the Flux Control Coefficients of a linear pathway with one or more feedback loops can be derived in a systematic manner by superimposing the effect of the feedback loop(s) on the expressions pertaining to the Flux Control Coefficients of the unregulated pathway.

Enzymes

Discrete-time random walks on diagrams (graphs) with cycles.

After a review of the diagram method for continuous-time random walks on graphs with cycles, the method is extended to discrete-time random walks. The basic theorems carry over formally from continuous time to discrete time. Three problems in tennis probabilities are used to illustrate random walks on discrete-time diagrams with cycles.

Humans

Learning directed acyclic graphs for ligands and receptors based on spatially resolved transcriptomic data of ovarian cancer.

To unravel the mechanism of immune activation and suppression within tumors, a critical step is to identify transcriptional signals governing cell-cell communication between tumor and immune/stromal cells in the tumor microenvironment. Central to this communication are interactions between secreted ligands and cell-surface receptors, creating a highly connected signaling network among cells. Recent advancements in in situ-omics profiling, particularly spatial transcriptomic (ST) technology, provide unique opportunities to directly characterize ligand-receptor signaling networks that power cell-cell communication. In this paper, we propose a novel statistical method, LRnetST, to characterize the ligand-receptor interaction networks between adjacent tumor and immune/stroma cells based on ST data. LRnetST utilizes a directed acyclic graph model with a novel approach to handle the zero-inflated distributions of ST data. It also leverages existing ligand-receptor regulation databases as prior information, and employs a bootstrap aggregation strategy to achieve robust network estimation. Application of LRnetST to ST data of high-grade serous ovarian tumor samples revealed both common and distinct ligand-receptor regulations across different tumors. Some of these interactions were validated through both a MERFISH dataset and a CosMx SMI dataset of independent ovarian tumor samples. These results cast light on biological processes relating to the communication between tumor and immune/stromal cells in ovarian tumors. An open-source R package of LRnetST is available on GitHub at https://github.com/jie108/LRnetST.

Humans

An extension of the graph theoretical approach to predict the secondary structure of large RNAs: the complex of 16S and 23S rRNAs from E. coli as a case study.

An algorithm using the graph theoretical approach to predict secondary structures of large nucleic acids is discussed. Reliability of prediction can be improved by incorporating available experimental data and sequence homology information. As a case study, this algorithm is applied to predict the secondary structure of the 16S-23S rRNA complex from E. coli. It was found that several structures of the complex can coexist. The computer program developed to predict the secondary structure of large RNAs can be run on IBM PC/AT compatible systems.

Algorithms