PubMed HealthSearch

SEARCH · PubMed Health

Results for “Graph”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Use of a rule based graph-theoretical system in evaluating the activity of a class of nucleoside analogues against human immunodeficiency virus.

A rule based graph-theoretical system has been used to evaluate qualitatively the activity of a class of nucleoside analogues against human immunodeficiency virus (HIV). The system identifies biologically relevant vertices (atoms) in the molecular graphs of the compounds which have the biological activity of interest. The idea is to relate biological activity with the structural or substructural characteristics of the compounds from the point of view of molecular topology (connectivity). The system brings vertices of similar or close topological environment in the respective compounds together and this is reflected in the ranges of values formed by a distance based index of the vertices, the 'distance exponent index (Dx)', where x is any real number. It is found that the system makes correct prediction of the activity of all the compounds (active as well as inactives) of both training set and the test set against HIV. It is also apparent from this study that the index D-4, which has been used here, can make a useful classification of the vertices according to their molecular environment and the system can produce significant result in a small as well as diverse data base.

Antiviral Agents

Descent graphs in pedigree analysis: applications to haplotyping, location scores, and marker-sharing statistics.

The introduction of stochastic methods in pedigree analysis has enabled geneticists to tackle computations intractable by standard deterministic methods. Until now these stochastic techniques have worked by running a Markov chain on the set of genetic descent states of a pedigree. Each descent state specifies the paths of gene flow in the pedigree and the founder alleles dropped down each path. The current paper follows up on a suggestion by Elizabeth Thompson that genetic descent graphs offer a more appropriate space for executing a Markov chain. A descent graph specifies the paths of gene flow but not the particular founder alleles traveling down the paths. This paper explores algorithms for implementing Thompson's suggestion for codominant markers in the context of automatic haplotyping, estimating location scores, and computing gene-clustering statistics for robust linkage analysis. Realistic numerical examples demonstrate the feasibility of the algorithms.

Alleles

A graph-grammar approach to represent causal, temporal and other contexts in an oncological patient record.

The data of a patient undergoing complex diagnostic and therapeutic procedures do not only form a simple chronology of events, but are closely related in many ways. Such data contexts include causal or temporal relationships, they express inconsistencies and revision processes, or describe patient-specific heuristics. The knowledge of data contexts supports the retrospective understanding of the medical decision-making process and is a valuable base for further treatment. Conventional data models usually neglect the problem of context knowledge, or simply use free text which is not processed by the program. In connection with the development of the knowledge-based system THEMPO (Therapy Management in Pediatric Oncology), which supports therapy and monitoring in pediatric oncology, a graph-grammar approach has been used to design and implement a graph-oriented patient model which allows the representation of non-trivial (causal, temporal, etc.) clinical contexts. For context acquisition a mouse-based tool has been developed allowing the physician to specify contexts in a comfortable graphical manner. Furthermore, the retrieval of contexts is realized with graphical tools as well.

Adverse Drug Reaction Reporting Systems

Application of attributed graphs in diagnostic pathology.

OBJECTIVE: To compute attributed graphs based upon calculation of the minimum spanning tree (MST) for various applications in diagnostic lung pathology. STUDY DESIGN: The study design included assistance in histologic diagnosis, confirmation of the diagnosis in single cases, measurement of texture alterations after induction chemotherapy, estimation of prognosis of operated-on lung cancer patients and analysis of lung cancer cells in association with differentiation markers. The histologic slides were Feulgen stained, and features of the integrated optical density (IOD) were associated with the nodes of the MST. The same procedure was applied to immunohistochemically and ligand histochemically stained slides by calculation of the mean staining intensity of the cytoplasm of tumor cells. A measure for structural entropy was introduced by computing the relative differences in distance and IOD between neighboring tumor cells in a 1/r2 field of force. In addition, the current of entropy was computed. RESULTS: Structural entropy reflects alterations in regular textures; the current of entropy is an especially good prognostic parameter in lung cancer. In immunochemistry and ligand histochemistry, construction of the attributed MST permits detailed insight into locally different staining behavior of tumor cells and immunocompetent cells. CONCLUSION: Attributed graphs contain important information that can be used for the estimation of survival or for confirmation of diagnostic entities, such as tumor cell types.

Adolescent

HallmarkGraph: a cancer hallmark informed graph neural network for classifying hierarchical tumor subtypes.

MOTIVATION: Accurate tumor subtype diagnosis is crucial for precision oncology, yet current methodologies face significant challenges. These include balancing model accuracy with interpretability and the high costs of generating multi-omics data in clinical settings. Moreover, there is a lack of validated models capable of classifying hierarchical tumor subtypes across a comprehensive pan-cancer cohort. RESULTS: We present a graph neural network, HallmarkGraph, the first biologically informed model developed to classify hierarchical tumor subtypes in human cancer. Inspired by cancer hallmarks, the model's architecture integrates transcriptome profiles and gene regulatory interactions to perform multi-label classification. We evaluate the model on a comprehensive pan-cancer cohort comprising 11 476 samples from 26 primary cancers with 405 subtypes up to eight levels. The model demonstrates exceptional performance, achieving 5-fold cross-validation accuracy between 85% and 99% for tumor subtypes labeled with increasing details of genomic information. It also shows good generalizability on a validation dataset of 887 samples, assessed using three metrics that consider tumor subtypes at individual, combined, and sample levels. Benchmarking and ablation experiments show that hallmark-based embeddings slightly influence model performance, while the integrated multilayer perceptron plays a significant role in determining classifier accuracy. Additionally, we use the SHAP method to link cancer hallmarks with genes, identifying key features that influence model decisions. Our findings present a biologically informed machine learning framework capable of tracking tumor transcriptomic trajectories and distinguishing inter- and intra-tumor heterogeneity in pan-cancer. This approach holds promise for enhancing cancer diagnostics. AVAILABILITY AND IMPLEMENTATION: HallmarkGraph is accessible at https://github.com/laixn/HallmarkGraph.

Humans

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

A General Framework for Branch Length Estimation in Ancestral Recombination Graphs.

Inference of Ancestral Recombination Graphs (ARGs) is of central interest in the analysis of genomic variation. ARGs can be specified in terms of topologies and coalescence times. The coalescence times are usually estimated using an informative prior derived from coalescent theory, but this may generate biased estimates and can also complicate downstream inferences based on ARGs. Here we introduce, POLEGON, a novel approach for estimating branch lengths for ARGs which uses an uninformative prior. Using extensive simulations, we show that this method provides improved estimates of coalescence times and lead to more accurate inferences of effective population sizes under a wide range of demographic assumptions (population expansion, bottleneck, split, etc). It also improves other downstream inferences including estimates of mutation rates. We apply the method to data from the 1000 Genomes Project to investigate population size histories and differential mutation signatures across populations. We also estimate coalescence times in the HLA region, and show that they exceed 30 million years in multiple segments.

Ancestral Recombination Graph

SPC: a SPectral Component approach leveraging Identity-by-Descent graphs to address recent population structure in genomic analysis.

Population structure is a well-known confounder in statistical genetics, particularly in genome-wide association studies (GWAS), where it can lead to inflated test statistics and spurious associations. Traditional methods, such as principal components (PCs), commonly used to adjust for population structure, are limited in capturing fine-scale, non-linear patterns that arise from recent demographic events - patterns that are crucial for understanding rare variant effects. To address this challenge, we propose a novel method called SPectral Components (SPCs), which leverages identity-by-descent (IBD) graphs to capture and transform local, non-linear fine-scale population structure into continuous representations that can be seamlessly integrated into genetic analysis pipelines. Using both simulated datasets and empirical data from the UK Biobank (N ≈ 420,000), we demonstrate that SPCs outperform PCs in adjusting for fine-scale population structure. In simulations, SPCs explained over 90% of the fine-scale population structure with fewer components, while PCs captured less than 5%. In the UK Biobank, SPCs reduced the inflation of p-values in the GWAS of an environmental-driven phenotype by 12% compared to PCs, while maintaining a similar performance to PCs in height, a highly heritable phenotype. Additionally, SPCs improved rare variant association analyses, reducing genomic inflation (e.g., from 7.6 to 1.2 in one analysis), and provided more accurate heritability estimates. Spatial autocorrelation analysis further confirmed the ability of SPCs to account for environmental effects, reducing Moran's I for both environmental and heritable phenotypes more effectively than PCs. Overall, our findings demonstrate that SPCs provide a robust, scalable adjustment for recent population structure, offering a powerful alternative or complement to PCs in large-scale biobank studies.

GWAS

A graph-based approach for the visualisation and analysis of bacterial pangenomes.

BACKGROUND: The advent of low cost, high throughput DNA sequencing has led to the availability of thousands of complete genome sequences for a wide variety of bacterial species. Examining and interpreting genetic variation on this scale represents a significant challenge to existing methods of data analysis and visualisation. RESULTS: Starting with the output of standard pangenome analysis tools, we describe the generation and analysis of interactive, 3D network graphs to explore the structure of bacterial populations, the distribution of genes across a population, and the syntenic order in which those genes occur, in the new open-source network analysis platform, Graphia. Both the analysis and the visualisation are scalable to datasets of thousands of genome sequences. CONCLUSIONS: We anticipate that the approaches presented here will be of great utility to the microbial research community, allowing faster, more intuitive, and flexible interaction with pangenome datasets, thereby enhancing interpretation of these complex data.

Bacteria

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans

A random graph model for the final-size distribution of household infections.

In epidemiological/disease control studies, one might be interested in estimating the parameters community probability infection (CPI) and the household secondary attack rate (SAR), as introduced by Longini and Koopman. The quasi-binomial distribution I (QBD I) with parameters n, p and theta, introduced by Consul, is proposed as a model for the final-size distribution of household infections, where p (CPI) is the probability of an individual being infected from the community and theta (SAR) is the rate of secondary transmission of infection within household. An individual can be infected either from within the household or from the community. Let X be the total number of infected members in a household of size n. Then the distribution of X is given by the QBD I with the probability mass function: (formula: see text) with 0 < p < 1, theta > or = 0 such that p + n theta < 1. The epidemic model is derived from a directed random graph. Data from influenza epidemics in Asian and American households are used to test the model and a comparison is made with the Longini-Koopman model. It is shown empirically that the QBD I is as good as the L-K model in describing the household infectious disease data, and both models provide almost identical estimates for community and household transmission parameters although they are derived from different perspectives and conditions.

Asia

A graph analysis of the relationship between population density and social pathology.

This article deals with systematic variables--population density and pathological effects--at the level of a community or organizational system. Existing studies of the relationship between population density and social pathology (e.g., mortality, mental illness) in humans have failed to determine a strictly causal relationship between density and pathology in a linear model. They have not established a direct link between density and pathology in a simple two-variable, linear causal approach. Eliminating other variables by statistical control has obscured the conditions of a complex phenomenon. By incorporating socioeconomic variables into a general system model, this paper suggests that the total configuration of social organization, adaptation, previous group experience, and environment determine the effects of population density. Using a method of graph analysis, a model is presented of the relationship between population density and social pathology which has a high degree of isomorphism with the empirical situation.

Adrenocorticotropic Hormone

Acyclic directed graphs for automatic image analysis of lung parenchymal geometry.

We have developed a simple method for encoding pulmonary alveolar geometry suitable for morphometric analysis and modeling. The approach involves the use of 8-bit video microscopic images of lung which are thresholded to produce binary images. Each alveolar wall was thinned to its midline to produce a skeletonized image. We present an algorithm for converting such an image to an acyclic directed graph. Such a representation is quite suitable for mathematical treatment, lending itself to problems involving simulation or modeling. We present as an example application, the measurement of dihedral angles in rat lung using conventional histological sections.

Algorithms

Uniqueness and multiplicity of steady states in monocyclic enzyme cascades: a graph-theoretic analysis.

Monocyclic enzyme cascades are important regulators of biochemical reactions in living cells. The reaction network may have either one steady state or many, depending on its structure. The occurrence of multiple steady states has important biological implications. A simple graph-theoretic method has been applied to five reaction mechanisms--which together cover many common monocyclic cascades--to determine which mechanisms generate just one steady state and which ones allow more than one state. It is shown that an unstable steady state in a multiplicity region can be usefully exploited and that in some cases transitions may occur between uniqueness and multiplicity regions. The possible effects of such transitions are discussed.

Animals

Graph-theoretical assignment of secondary structure in multidimensional protein NMR spectra: application to the lac repressor headpiece.

A novel procedure is presented for the automatic identification of secondary structures in proteins from their corresponding NOE data. The method uses a branch of mathematics known as graph theory to identify prescribed NOE connectivity patterns characteristic of the regular secondary structures. Resonance assignment is achieved by connecting these patterns of secondary structure together, thereby matching the connected spin systems to specific segments of the protein sequence. The method known as SERENDIPITY refers to a set of routines developed in a modular fashion, where each program has one or several well-defined tasks. NOE templates for several secondary structure motifs have been developed and the method has been successfully applied to data obtained from NOESY-type spectra. The present report describes the application of the SERENDIPITY protocol to a 3D NOESY-HMQC spectrum of the 15N-labelled lac repressor headpiece protein. The application demonstrates that, under favourable conditions, fully automated identification of secondary structures and semi-automated assignment are feasible.

Amino Acid Sequence

The graph of loudness recruitment (ABLB-test).

The ABLB-test compares the subjective loudness in patients with unilateral sensorineural hearing loss. The graph of this comparison of loudness at different suprathreshold intensities is of interest for understanding the function of the hair cells.

Auditory Threshold

Molecular fragments associated with non-genotoxic carcinogens, as detected using a software program based on graph theory: their usefulness to predict carcinogenicity.

We assembled 390 chemicals with a structure non-alerting to DNA-reactivity (145 carcinogens and 245 non-carcinogens) for which rodent carcinogenicity data were available. These non-alerting chemicals were defined by the absence in their molecules of DNA-reactive (directly or after metabolic activation) alerting structures, as described by Ashby and coworkers (Mutat. Res., 204 (1988) 17-115; Mutat. Res., 223 (1989) 73-103; Mutat. Res., 257 (1991) 209-227; Mutat. Res., 286 (1993) 3-74). Using our software program based on graph theory we analyzed the compounds in order to estimate the program's ability to predict nonalerting carcinogens. Our software fragmented the structural formula of the chemicals into all possible fragments of contiguous atoms with size between 2 and 8 (non-hydrogen) atoms and learned about statistically significant fragments from a training set of chemicals. These fragments were used to predict carcinogenicity or lack thereof in a verification set of compounds. For 390 runs of the software program we used (n - 1) of the chemicals as a training set, to predict the excluded chemical at each run (as a test set). Using two different probability thresholds to select significant fragments (P = 0.05 and P = 0.125 1-tailed according to binomial distribution), we performed two analyses: in the better one (P = 0.05) 19% of the molecules tested lacked significant fragments, for the remaining 81% the observed level of accuracy of the prediction was 66.0% against an expected level of accuracy of 51.7%. The difference was highly significant (P < 0.0001). We also examined the more significant activating fragments (biophores) and discussed at length both their biological plausibility and the working hypothesis that additional alerting structures for carcinogenicity (not only those related to genotoxicity) can be detected using this type of SAR approach. This new class of alerting structures could identify subfamilies of congeneric analogs active through mechanisms of receptor mediated carcinogenesis.

Carcinogens

Bond graphs and the exploitation of power conserving transformations.

Use is made of the guaranteed energy conservation property of any bond graph model (presuming that consistent energy variables are used). Thus power transformations are power conserving, and this property is exploited with respect to multiport transformers where many effort and flow variables may be involved. In these cases the relationship among the effort (flow) variables across the transformer may be easier to derive than the relationship among the flows (efforts). The power conserving nature of the transformer permits immediate derivation of the alternate variable relationship. This formulation procedure is applied to the reflex reaction of the arm.

Energy Metabolism