PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Graph”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

[Neighborhood graphs and image processing. Contribution to images of immunohistochemical staining].

Devising an image analyzer dedicated to the automatic quantification of immunohistochemical staining for clinical oncology implies developing a method for the delimitation of tumoral cell nests, setting aside tumoral stroma, while accounting for the topology of the staining. The representation of images by neighborhood graphs can bring an answer to both requirements. In this paper, a methodological approach is presented. It consists in a preliminary study dealing with nuclear immunostaining images of breast cancer. Segmentation of the graph structure allows to separate clusters of cancer cells and the analysis of this structure can account for the focal or diffuse aspect of the staining within the tumor.

Breast Neoplasms↗

Formal classification of medical concept descriptions: graph-oriented operators.

A crucial component of a medical concept representation system is the classifier. It requires features that are not sufficiently supported by current logic based formalisms like description logics and conceptual graphs. Those features are, for instance, the representation of partitive and spatial relations and their impact on subsumption. This paper introduces graph oriented classification operators for a concept representation language with normal forms. Emphasis is on the separation of generic and partitive relations and on the mutual interdependence of subsumption and part-whole. For that purpose operators are given for formal subsumption, formal part-whole, subsumptive part-whole and part-sensitive subsumption. These operators are based on the formal structure of concept descriptions and on explicitly introduced generic and partitive relationships between their constituents.

Abstracting and Indexing↗

HallmarkGraph: a cancer hallmark informed graph neural network for classifying hierarchical tumor subtypes.

MOTIVATION: Accurate tumor subtype diagnosis is crucial for precision oncology, yet current methodologies face significant challenges. These include balancing model accuracy with interpretability and the high costs of generating multi-omics data in clinical settings. Moreover, there is a lack of validated models capable of classifying hierarchical tumor subtypes across a comprehensive pan-cancer cohort. RESULTS: We present a graph neural network, HallmarkGraph, the first biologically informed model developed to classify hierarchical tumor subtypes in human cancer. Inspired by cancer hallmarks, the model's architecture integrates transcriptome profiles and gene regulatory interactions to perform multi-label classification. We evaluate the model on a comprehensive pan-cancer cohort comprising 11 476 samples from 26 primary cancers with 405 subtypes up to eight levels. The model demonstrates exceptional performance, achieving 5-fold cross-validation accuracy between 85% and 99% for tumor subtypes labeled with increasing details of genomic information. It also shows good generalizability on a validation dataset of 887 samples, assessed using three metrics that consider tumor subtypes at individual, combined, and sample levels. Benchmarking and ablation experiments show that hallmark-based embeddings slightly influence model performance, while the integrated multilayer perceptron plays a significant role in determining classifier accuracy. Additionally, we use the SHAP method to link cancer hallmarks with genes, identifying key features that influence model decisions. Our findings present a biologically informed machine learning framework capable of tracking tumor transcriptomic trajectories and distinguishing inter- and intra-tumor heterogeneity in pan-cancer. This approach holds promise for enhancing cancer diagnostics. AVAILABILITY AND IMPLEMENTATION: HallmarkGraph is accessible at https://github.com/laixn/HallmarkGraph.

Humans↗

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks↗

A General Framework for Branch Length Estimation in Ancestral Recombination Graphs.

Inference of Ancestral Recombination Graphs (ARGs) is of central interest in the analysis of genomic variation. ARGs can be specified in terms of topologies and coalescence times. The coalescence times are usually estimated using an informative prior derived from coalescent theory, but this may generate biased estimates and can also complicate downstream inferences based on ARGs. Here we introduce, POLEGON, a novel approach for estimating branch lengths for ARGs which uses an uninformative prior. Using extensive simulations, we show that this method provides improved estimates of coalescence times and lead to more accurate inferences of effective population sizes under a wide range of demographic assumptions (population expansion, bottleneck, split, etc). It also improves other downstream inferences including estimates of mutation rates. We apply the method to data from the 1000 Genomes Project to investigate population size histories and differential mutation signatures across populations. We also estimate coalescence times in the HLA region, and show that they exceed 30 million years in multiple segments.

Ancestral Recombination Graph↗

SPC: a SPectral Component approach leveraging Identity-by-Descent graphs to address recent population structure in genomic analysis.

Population structure is a well-known confounder in statistical genetics, particularly in genome-wide association studies (GWAS), where it can lead to inflated test statistics and spurious associations. Traditional methods, such as principal components (PCs), commonly used to adjust for population structure, are limited in capturing fine-scale, non-linear patterns that arise from recent demographic events - patterns that are crucial for understanding rare variant effects. To address this challenge, we propose a novel method called SPectral Components (SPCs), which leverages identity-by-descent (IBD) graphs to capture and transform local, non-linear fine-scale population structure into continuous representations that can be seamlessly integrated into genetic analysis pipelines. Using both simulated datasets and empirical data from the UK Biobank (N ≈ 420,000), we demonstrate that SPCs outperform PCs in adjusting for fine-scale population structure. In simulations, SPCs explained over 90% of the fine-scale population structure with fewer components, while PCs captured less than 5%. In the UK Biobank, SPCs reduced the inflation of p-values in the GWAS of an environmental-driven phenotype by 12% compared to PCs, while maintaining a similar performance to PCs in height, a highly heritable phenotype. Additionally, SPCs improved rare variant association analyses, reducing genomic inflation (e.g., from 7.6 to 1.2 in one analysis), and provided more accurate heritability estimates. Spatial autocorrelation analysis further confirmed the ability of SPCs to account for environmental effects, reducing Moran's I for both environmental and heritable phenotypes more effectively than PCs. Overall, our findings demonstrate that SPCs provide a robust, scalable adjustment for recent population structure, offering a powerful alternative or complement to PCs in large-scale biobank studies.

GWAS↗

A graph-based approach for the visualisation and analysis of bacterial pangenomes.

BACKGROUND: The advent of low cost, high throughput DNA sequencing has led to the availability of thousands of complete genome sequences for a wide variety of bacterial species. Examining and interpreting genetic variation on this scale represents a significant challenge to existing methods of data analysis and visualisation. RESULTS: Starting with the output of standard pangenome analysis tools, we describe the generation and analysis of interactive, 3D network graphs to explore the structure of bacterial populations, the distribution of genes across a population, and the syntenic order in which those genes occur, in the new open-source network analysis platform, Graphia. Both the analysis and the visualisation are scalable to datasets of thousands of genome sequences. CONCLUSIONS: We anticipate that the approaches presented here will be of great utility to the microbial research community, allowing faster, more intuitive, and flexible interaction with pangenome datasets, thereby enhancing interpretation of these complex data.

Bacteria↗

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans↗

Efficacy of extended-release niacin with lovastatin for hypercholesterolemia: assessing all reasonable doses with innovative surface graph analysis.

BACKGROUND: Combination therapy to improve the total lipid profile may achieve greater coronary risk reductions than lowering low-density lipoprotein cholesterol (LDL-C) alone. A new extended-release niacin (niacin ER)/lovastatin tablet substantially lowers LDL-C, triglyceride, and lipoprotein(a) levels and raises high-density lipoprotein cholesterol (HDL-C) level. We evaluated these serum lipid responses to niacin ER/lovastatin at all clinically reasonable doses. METHODS: Men (n = 85) and women (n = 79) with type IIa or IIb primary hyperlipidemia after diet were randomized among 5 parallel treatment arms. Each arm had 5 sequential 4-week treatment periods: niacin ER (starting at 500 mg/d, increasing in 500-mg increments to 2500 mg/d); lovastatin (starting at 10 mg, increasing to 20 mg, then 40 mg/d); and 3 combinations arms, each with a constant lovastatin dose and escalating niacin ER doses. RESULTS: For primary comparisons, mean LDL-C level reductions from baseline were greater with niacin ER/lovastatin (1500/20 mg) than with lovastatin (20 mg) (35% vs 22%, P<.001) and with niacin ER/lovastatin (2000/40 mg) than with lovastatin (40 mg) (46% vs 24%, P<.001). Each 500-mg increase in niacin ER, on average, decreased LDL-C levels an additional 4% and increased HDL-C levels 8%. The maximum recommended dose (2000/40 mg/d) increased HDL-C levels 29% and decreased LDL-C levels 46%, triglyceride levels 38%, and lipoprotein(a) levels 14%. All lipid responses were dose dependent and generally additive. Graphs of the dose-response relationships as 3-dimensional surfaces documented the strength and consistency of these responses. CONCLUSIONS: Niacin ER/lovastatin combination therapy substantially improves 4 major lipoprotein levels associated with atherosclerotic disease. Dose-response surfaces provide a practical guide for dose selection.

Delayed-Action Preparations↗

Handling context-sensitivity in protein structures using graph theory: bona fide prediction.

We constructed five comparative models in a blind manner for the second meeting on the Critical Assessment of protein Structure Prediction methods (CASP2). The method used is based on a novel graph-theoretic clique-finding approach, and attempts to address the problem of interconnected structural changes in the comparative modeling of protein structures. We discuss briefly how the method is used for protein structure prediction, and detail how it performs in the blind tests. We find that compared to CASP1, significant improvements in building insertions and deletions and sidechain conformations have been achieved.

Amino Acid Sequence↗

RMS/coverage graphs: a qualitative method for comparing three-dimensional protein structure predictions.

Evaluating a set of protein structure predictions is difficult as each prediction may omit different residues and different parts of the structure may have different accuracies. A method is described that captures the best results from a large number of alternative sequence-dependent structural superpositions between a prediction and the experimental structure and represents them as a single line on a graph. Applied to CASP2 and CASP3 data the best predictions stand out visually in most cases, as judged by manual inspection. The results from this method applied to CASP data are available from the URLs http:/(/)PredictionCenter. llnl.gov/casp3/results/th/ and http:/(/)www.sanger.ac.uk/ approximately th/casp/.

Computer Graphics↗

A random graph model for the final-size distribution of household infections.

In epidemiological/disease control studies, one might be interested in estimating the parameters community probability infection (CPI) and the household secondary attack rate (SAR), as introduced by Longini and Koopman. The quasi-binomial distribution I (QBD I) with parameters n, p and theta, introduced by Consul, is proposed as a model for the final-size distribution of household infections, where p (CPI) is the probability of an individual being infected from the community and theta (SAR) is the rate of secondary transmission of infection within household. An individual can be infected either from within the household or from the community. Let X be the total number of infected members in a household of size n. Then the distribution of X is given by the QBD I with the probability mass function: (formula: see text) with 0 < p < 1, theta > or = 0 such that p + n theta < 1. The epidemic model is derived from a directed random graph. Data from influenza epidemics in Asian and American households are used to test the model and a comparison is made with the Longini-Koopman model. It is shown empirically that the QBD I is as good as the L-K model in describing the household infectious disease data, and both models provide almost identical estimates for community and household transmission parameters although they are derived from different perspectives and conditions.

Asia↗

Clusters in alpha/beta barrel proteins: implications for protein structure, function, and folding: a graph theoretical approach.

The alpha/beta barrel fold is adopted by most enzymes performing a variety of catalytic reactions, but with very low sequence similarity. In order to understand the stabilizing interactions important in maintaining the alpha/beta barrel fold, we have identified residue clusters in a dataset of 36 alpha/beta barrel proteins that have less than 10% sequence identity within themselves. A graph theoretical algorithm is used to identify backbone clusters. This approach uses the global information of the nonbonded interaction in the alpha/beta barrel fold for the clustering procedure. The nonbonded interactions are represented mathematically in the form of an adjacency matrix. On diagonalizing the adjacency matrix, clusters and cluster centers are obtained from the highest eigenvalue and its corresponding vector components. Residue clusters are identified in the strand regions forming the beta barrel and are topologically conserved in all 36 proteins studied. The residues forming the cluster in each of the alpha/beta protein are also conserved among the sequences belonging to the same family. The cluster centers are found to occur in the middle of the strands or in the C-terminal of the strands. In most cases, the residues forming the clusters are part of the active site or are located close to the active site. The folding nucleus of the alpha/beta fold is predicted based on hydrophobicity index evaluation of residues and identification of cluster centers. The predicted nucleation sites are found to occur mostly in the middle of the strands. Proteins 2001;43:103-112.

Algorithms↗

A graph analysis of the relationship between population density and social pathology.

This article deals with systematic variables--population density and pathological effects--at the level of a community or organizational system. Existing studies of the relationship between population density and social pathology (e.g., mortality, mental illness) in humans have failed to determine a strictly causal relationship between density and pathology in a linear model. They have not established a direct link between density and pathology in a simple two-variable, linear causal approach. Eliminating other variables by statistical control has obscured the conditions of a complex phenomenon. By incorporating socioeconomic variables into a general system model, this paper suggests that the total configuration of social organization, adaptation, previous group experience, and environment determine the effects of population density. Using a method of graph analysis, a model is presented of the relationship between population density and social pathology which has a high degree of isomorphism with the empirical situation.

Adrenocorticotropic Hormone↗

BRIDGE (building the relationship between body image and disordered eating graph and explanation): a tool for parents and professionals.

BRIDGE (Building the Relationship Between Body Image and Disordered Eating Graph and Explanation) is a tool for mental health professionals, parents, teachers, and the general public. BRIDGE offers the potential for a new level of understanding between body image and disordered eating behaviors. This tool provides a framework that describes the connection between the attitudes we hold about our bodies and the corresponding behaviors we may practice. This practical tool takes fundamental concepts established in the eating disorder literature and presents them in a basic and easily understood manner. This paper invites further testing of this educational tool for parents, teachers, and professionals in a number of contexts including prevention and treatment programs.

Journal Article↗

Calculation of flip angles for echo trains with predefined amplitudes with the extended phase graph (EPG)-algorithm: principles and applications to hyperecho and TRAPS sequences.

The article presents an algorithm for calculation of flip angles in multiecho experiments to generate echoes with predefined amplitudes based on the extended phase graph algorithm. The algorithm can be used to optimize the echo envelope and thus the point spread function (PSF) in hyperecho and TRAPS (transition into the pseudosteady state) experiments while minimizing the total RF power. Implementations at 3 T using echo trains with Gaussian and Lorentzian PSF demonstrate a reduction in RF power by a factor of 3-5 while maintaining high image quality.

Algorithms↗

Protein flexibility predictions using graph theory.

Techniques from graph theory are applied to analyze the bond networks in proteins and identify the flexible and rigid regions. The bond network consists of distance constraints defined by the covalent and hydrogen bonds and salt bridges in the protein, identified by geometric and energetic criteria. We use an algorithm that counts the degrees of freedom within this constraint network and that identifies all the rigid and flexible substructures in the protein, including overconstrained regions (with more crosslinking bonds than are needed to rigidify the region) and underconstrained or flexible regions, in which dihedral bond rotations can occur. The number of extra constraints or remaining degrees of bond-rotational freedom within a substructure quantifies its relative rigidity/flexibility and provides a flexibility index for each bond in the structure. This novel computational procedure, first used in the analysis of glassy materials, is approximately a million times faster than molecular dynamics simulations and captures the essential conformational flexibility of the protein main and side-chains from analysis of a single, static three-dimensional structure. This approach is demonstrated by comparison with experimental measures of flexibility for three proteins in which hinge and loop motion are essential for biological function: HIV protease, adenylate kinase, and dihydrofolate reductase.

Adenylate Kinase↗

Large-scale prediction of disulphide bridges using kernel methods, two-dimensional recursive neural networks, and weighted graph matching.

The formation of disulphide bridges between cysteines plays an important role in protein folding, structure, function, and evolution. Here, we develop new methods for predicting disulphide bridges in proteins. We first build a large curated data set of proteins containing disulphide bridges to extract relevant statistics. We then use kernel methods to predict whether a given protein chain contains intrachain disulphide bridges or not, and recursive neural networks to predict the bonding probabilities of each pair of cysteines in the chain. These probabilities in turn lead to an accurate estimation of the total number of disulphide bridges and to a weighted graph matching problem that can be addressed efficiently to infer the global disulphide bridge connectivity pattern. This approach can be applied both in situations where the bonded state of each cysteine is known, or in ab initio mode where the state is unknown. Furthermore, it can easily cope with chains containing an arbitrary number of disulphide bridges, overcoming one of the major limitations of previous approaches. It can classify individual cysteine residues as bonded or nonbonded with 87% specificity and 89% sensitivity. The estimate for the total number of bridges in each chain is correct 71% of the times, and within one from the true value over 94% of the times. The prediction of the overall disulphide connectivity pattern is exact in about 51% of the chains. In addition to using profiles in the input to leverage evolutionary information, including true (but not predicted) secondary structure and solvent accessibility information yields small but noticeable improvements. Finally, once the system is trained, predictions can be computed rapidly on a proteomic or protein-engineering scale. The disulphide bridge prediction server (DIpro), software, and datasets are available through www.igb.uci.edu/servers/psss.html.

Amino Acid Sequence↗