PubMed HealthSearch

SEARCH · PubMed Health

Results for “Graph”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Application of attributed graphs in diagnostic pathology.

OBJECTIVE: To compute attributed graphs based upon calculation of the minimum spanning tree (MST) for various applications in diagnostic lung pathology. STUDY DESIGN: The study design included assistance in histologic diagnosis, confirmation of the diagnosis in single cases, measurement of texture alterations after induction chemotherapy, estimation of prognosis of operated-on lung cancer patients and analysis of lung cancer cells in association with differentiation markers. The histologic slides were Feulgen stained, and features of the integrated optical density (IOD) were associated with the nodes of the MST. The same procedure was applied to immunohistochemically and ligand histochemically stained slides by calculation of the mean staining intensity of the cytoplasm of tumor cells. A measure for structural entropy was introduced by computing the relative differences in distance and IOD between neighboring tumor cells in a 1/r2 field of force. In addition, the current of entropy was computed. RESULTS: Structural entropy reflects alterations in regular textures; the current of entropy is an especially good prognostic parameter in lung cancer. In immunochemistry and ligand histochemistry, construction of the attributed MST permits detailed insight into locally different staining behavior of tumor cells and immunocompetent cells. CONCLUSION: Attributed graphs contain important information that can be used for the estimation of survival or for confirmation of diagnostic entities, such as tumor cell types.

Adolescent

[Neighborhood graphs and image processing. Contribution to images of immunohistochemical staining].

Devising an image analyzer dedicated to the automatic quantification of immunohistochemical staining for clinical oncology implies developing a method for the delimitation of tumoral cell nests, setting aside tumoral stroma, while accounting for the topology of the staining. The representation of images by neighborhood graphs can bring an answer to both requirements. In this paper, a methodological approach is presented. It consists in a preliminary study dealing with nuclear immunostaining images of breast cancer. Segmentation of the graph structure allows to separate clusters of cancer cells and the analysis of this structure can account for the focal or diffuse aspect of the staining within the tumor.

Breast Neoplasms

Formal classification of medical concept descriptions: graph-oriented operators.

A crucial component of a medical concept representation system is the classifier. It requires features that are not sufficiently supported by current logic based formalisms like description logics and conceptual graphs. Those features are, for instance, the representation of partitive and spatial relations and their impact on subsumption. This paper introduces graph oriented classification operators for a concept representation language with normal forms. Emphasis is on the separation of generic and partitive relations and on the mutual interdependence of subsumption and part-whole. For that purpose operators are given for formal subsumption, formal part-whole, subsumptive part-whole and part-sensitive subsumption. These operators are based on the formal structure of concept descriptions and on explicitly introduced generic and partitive relationships between their constituents.

Abstracting and Indexing

HallmarkGraph: a cancer hallmark informed graph neural network for classifying hierarchical tumor subtypes.

MOTIVATION: Accurate tumor subtype diagnosis is crucial for precision oncology, yet current methodologies face significant challenges. These include balancing model accuracy with interpretability and the high costs of generating multi-omics data in clinical settings. Moreover, there is a lack of validated models capable of classifying hierarchical tumor subtypes across a comprehensive pan-cancer cohort. RESULTS: We present a graph neural network, HallmarkGraph, the first biologically informed model developed to classify hierarchical tumor subtypes in human cancer. Inspired by cancer hallmarks, the model's architecture integrates transcriptome profiles and gene regulatory interactions to perform multi-label classification. We evaluate the model on a comprehensive pan-cancer cohort comprising 11 476 samples from 26 primary cancers with 405 subtypes up to eight levels. The model demonstrates exceptional performance, achieving 5-fold cross-validation accuracy between 85% and 99% for tumor subtypes labeled with increasing details of genomic information. It also shows good generalizability on a validation dataset of 887 samples, assessed using three metrics that consider tumor subtypes at individual, combined, and sample levels. Benchmarking and ablation experiments show that hallmark-based embeddings slightly influence model performance, while the integrated multilayer perceptron plays a significant role in determining classifier accuracy. Additionally, we use the SHAP method to link cancer hallmarks with genes, identifying key features that influence model decisions. Our findings present a biologically informed machine learning framework capable of tracking tumor transcriptomic trajectories and distinguishing inter- and intra-tumor heterogeneity in pan-cancer. This approach holds promise for enhancing cancer diagnostics. AVAILABILITY AND IMPLEMENTATION: HallmarkGraph is accessible at https://github.com/laixn/HallmarkGraph.

Humans

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

A General Framework for Branch Length Estimation in Ancestral Recombination Graphs.

Inference of Ancestral Recombination Graphs (ARGs) is of central interest in the analysis of genomic variation. ARGs can be specified in terms of topologies and coalescence times. The coalescence times are usually estimated using an informative prior derived from coalescent theory, but this may generate biased estimates and can also complicate downstream inferences based on ARGs. Here we introduce, POLEGON, a novel approach for estimating branch lengths for ARGs which uses an uninformative prior. Using extensive simulations, we show that this method provides improved estimates of coalescence times and lead to more accurate inferences of effective population sizes under a wide range of demographic assumptions (population expansion, bottleneck, split, etc). It also improves other downstream inferences including estimates of mutation rates. We apply the method to data from the 1000 Genomes Project to investigate population size histories and differential mutation signatures across populations. We also estimate coalescence times in the HLA region, and show that they exceed 30 million years in multiple segments.

Ancestral Recombination Graph

SPC: a SPectral Component approach leveraging Identity-by-Descent graphs to address recent population structure in genomic analysis.

Population structure is a well-known confounder in statistical genetics, particularly in genome-wide association studies (GWAS), where it can lead to inflated test statistics and spurious associations. Traditional methods, such as principal components (PCs), commonly used to adjust for population structure, are limited in capturing fine-scale, non-linear patterns that arise from recent demographic events - patterns that are crucial for understanding rare variant effects. To address this challenge, we propose a novel method called SPectral Components (SPCs), which leverages identity-by-descent (IBD) graphs to capture and transform local, non-linear fine-scale population structure into continuous representations that can be seamlessly integrated into genetic analysis pipelines. Using both simulated datasets and empirical data from the UK Biobank (N ≈ 420,000), we demonstrate that SPCs outperform PCs in adjusting for fine-scale population structure. In simulations, SPCs explained over 90% of the fine-scale population structure with fewer components, while PCs captured less than 5%. In the UK Biobank, SPCs reduced the inflation of p-values in the GWAS of an environmental-driven phenotype by 12% compared to PCs, while maintaining a similar performance to PCs in height, a highly heritable phenotype. Additionally, SPCs improved rare variant association analyses, reducing genomic inflation (e.g., from 7.6 to 1.2 in one analysis), and provided more accurate heritability estimates. Spatial autocorrelation analysis further confirmed the ability of SPCs to account for environmental effects, reducing Moran's I for both environmental and heritable phenotypes more effectively than PCs. Overall, our findings demonstrate that SPCs provide a robust, scalable adjustment for recent population structure, offering a powerful alternative or complement to PCs in large-scale biobank studies.

GWAS

A graph-based approach for the visualisation and analysis of bacterial pangenomes.

BACKGROUND: The advent of low cost, high throughput DNA sequencing has led to the availability of thousands of complete genome sequences for a wide variety of bacterial species. Examining and interpreting genetic variation on this scale represents a significant challenge to existing methods of data analysis and visualisation. RESULTS: Starting with the output of standard pangenome analysis tools, we describe the generation and analysis of interactive, 3D network graphs to explore the structure of bacterial populations, the distribution of genes across a population, and the syntenic order in which those genes occur, in the new open-source network analysis platform, Graphia. Both the analysis and the visualisation are scalable to datasets of thousands of genome sequences. CONCLUSIONS: We anticipate that the approaches presented here will be of great utility to the microbial research community, allowing faster, more intuitive, and flexible interaction with pangenome datasets, thereby enhancing interpretation of these complex data.

Bacteria

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans

Handling context-sensitivity in protein structures using graph theory: bona fide prediction.

We constructed five comparative models in a blind manner for the second meeting on the Critical Assessment of protein Structure Prediction methods (CASP2). The method used is based on a novel graph-theoretic clique-finding approach, and attempts to address the problem of interconnected structural changes in the comparative modeling of protein structures. We discuss briefly how the method is used for protein structure prediction, and detail how it performs in the blind tests. We find that compared to CASP1, significant improvements in building insertions and deletions and sidechain conformations have been achieved.

Amino Acid Sequence

A random graph model for the final-size distribution of household infections.

In epidemiological/disease control studies, one might be interested in estimating the parameters community probability infection (CPI) and the household secondary attack rate (SAR), as introduced by Longini and Koopman. The quasi-binomial distribution I (QBD I) with parameters n, p and theta, introduced by Consul, is proposed as a model for the final-size distribution of household infections, where p (CPI) is the probability of an individual being infected from the community and theta (SAR) is the rate of secondary transmission of infection within household. An individual can be infected either from within the household or from the community. Let X be the total number of infected members in a household of size n. Then the distribution of X is given by the QBD I with the probability mass function: (formula: see text) with 0 < p < 1, theta > or = 0 such that p + n theta < 1. The epidemic model is derived from a directed random graph. Data from influenza epidemics in Asian and American households are used to test the model and a comparison is made with the Longini-Koopman model. It is shown empirically that the QBD I is as good as the L-K model in describing the household infectious disease data, and both models provide almost identical estimates for community and household transmission parameters although they are derived from different perspectives and conditions.

Asia

A graph analysis of the relationship between population density and social pathology.

This article deals with systematic variables--population density and pathological effects--at the level of a community or organizational system. Existing studies of the relationship between population density and social pathology (e.g., mortality, mental illness) in humans have failed to determine a strictly causal relationship between density and pathology in a linear model. They have not established a direct link between density and pathology in a simple two-variable, linear causal approach. Eliminating other variables by statistical control has obscured the conditions of a complex phenomenon. By incorporating socioeconomic variables into a general system model, this paper suggests that the total configuration of social organization, adaptation, previous group experience, and environment determine the effects of population density. Using a method of graph analysis, a model is presented of the relationship between population density and social pathology which has a high degree of isomorphism with the empirical situation.

Adrenocorticotropic Hormone

Acyclic directed graphs for automatic image analysis of lung parenchymal geometry.

We have developed a simple method for encoding pulmonary alveolar geometry suitable for morphometric analysis and modeling. The approach involves the use of 8-bit video microscopic images of lung which are thresholded to produce binary images. Each alveolar wall was thinned to its midline to produce a skeletonized image. We present an algorithm for converting such an image to an acyclic directed graph. Such a representation is quite suitable for mathematical treatment, lending itself to problems involving simulation or modeling. We present as an example application, the measurement of dihedral angles in rat lung using conventional histological sections.

Algorithms

Uniqueness and multiplicity of steady states in monocyclic enzyme cascades: a graph-theoretic analysis.

Monocyclic enzyme cascades are important regulators of biochemical reactions in living cells. The reaction network may have either one steady state or many, depending on its structure. The occurrence of multiple steady states has important biological implications. A simple graph-theoretic method has been applied to five reaction mechanisms--which together cover many common monocyclic cascades--to determine which mechanisms generate just one steady state and which ones allow more than one state. It is shown that an unstable steady state in a multiplicity region can be usefully exploited and that in some cases transitions may occur between uniqueness and multiplicity regions. The possible effects of such transitions are discussed.

Animals

Cytological, genetic and evolutionary functions of chiasmata based on chiasma graph analysis.

The nature of the chiasma as a cytological parameter for analysing cross-over was reexamined quantitatively by an improved chiasma graph method. It was reconfirmed in Mus platythrix (n =13) that interstitial chiasmata at diakinesis are distributed randomly and almost uniformly along bivalents except for the centromere and telomere regions. The size of these chiasma blank regions was consistently 0.8% of the total length of haploid autosomes in all chromosomes. There was a minimum value of chiasma interference distance between two adjacent chiasmata, which was constantly 1.8% in all chromosomes. The chiasma frequency at diakinesis was 20.1+/-2. 0 by the conventional method including terminal chiasmata. However, the primed in situ labeling technique revealed that terminal chiasmata were mostly telomere-telomere associations. From these data and also from recent molecular data we concluded that the terminal chiasma is cytologically functional for ensuring the normal disjunction of bivalents at anaphase I, but genetically non-functional for shuffling genes. The chiasma frequency excluding terminal chiasmata was 14.6+/-1.8. Reexamination of the chiasma frequency of 106 animal species revealed that the chiasma frequency increased linearly in proportion to the haploid chromosome number in spite of remarkable difference in their genome size. The increase in chiasma frequency would be evolution-adaptive, because gene shuffling is expected to be accelerated in species with high chromosome numbers.

Animals

Graph-theoretical assignment of secondary structure in multidimensional protein NMR spectra: application to the lac repressor headpiece.

A novel procedure is presented for the automatic identification of secondary structures in proteins from their corresponding NOE data. The method uses a branch of mathematics known as graph theory to identify prescribed NOE connectivity patterns characteristic of the regular secondary structures. Resonance assignment is achieved by connecting these patterns of secondary structure together, thereby matching the connected spin systems to specific segments of the protein sequence. The method known as SERENDIPITY refers to a set of routines developed in a modular fashion, where each program has one or several well-defined tasks. NOE templates for several secondary structure motifs have been developed and the method has been successfully applied to data obtained from NOESY-type spectra. The present report describes the application of the SERENDIPITY protocol to a 3D NOESY-HMQC spectrum of the 15N-labelled lac repressor headpiece protein. The application demonstrates that, under favourable conditions, fully automated identification of secondary structures and semi-automated assignment are feasible.

Amino Acid Sequence

The graph of loudness recruitment (ABLB-test).

The ABLB-test compares the subjective loudness in patients with unilateral sensorineural hearing loss. The graph of this comparison of loudness at different suprathreshold intensities is of interest for understanding the function of the hair cells.

Auditory Threshold

Molecular fragments associated with non-genotoxic carcinogens, as detected using a software program based on graph theory: their usefulness to predict carcinogenicity.

We assembled 390 chemicals with a structure non-alerting to DNA-reactivity (145 carcinogens and 245 non-carcinogens) for which rodent carcinogenicity data were available. These non-alerting chemicals were defined by the absence in their molecules of DNA-reactive (directly or after metabolic activation) alerting structures, as described by Ashby and coworkers (Mutat. Res., 204 (1988) 17-115; Mutat. Res., 223 (1989) 73-103; Mutat. Res., 257 (1991) 209-227; Mutat. Res., 286 (1993) 3-74). Using our software program based on graph theory we analyzed the compounds in order to estimate the program's ability to predict nonalerting carcinogens. Our software fragmented the structural formula of the chemicals into all possible fragments of contiguous atoms with size between 2 and 8 (non-hydrogen) atoms and learned about statistically significant fragments from a training set of chemicals. These fragments were used to predict carcinogenicity or lack thereof in a verification set of compounds. For 390 runs of the software program we used (n - 1) of the chemicals as a training set, to predict the excluded chemical at each run (as a test set). Using two different probability thresholds to select significant fragments (P = 0.05 and P = 0.125 1-tailed according to binomial distribution), we performed two analyses: in the better one (P = 0.05) 19% of the molecules tested lacked significant fragments, for the remaining 81% the observed level of accuracy of the prediction was 66.0% against an expected level of accuracy of 51.7%. The difference was highly significant (P < 0.0001). We also examined the more significant activating fragments (biophores) and discussed at length both their biological plausibility and the working hypothesis that additional alerting structures for carcinogenicity (not only those related to genotoxicity) can be detected using this type of SAR approach. This new class of alerting structures could identify subfamilies of congeneric analogs active through mechanisms of receptor mediated carcinogenesis.

Carcinogens