PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Graph genome”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

VirBinn improves viral genome binning from metagenomic Hi-C through graph diffusion.

MOTIVATION: Metagenomic Hi-C provides in situ proximity signals that can improve genome binning and enable virus-host-association analysis. However, viral genome recovery remains difficult because virus-virus Hi-C contact matrices are extremely sparse. Viral genomes are small, often low-abundance, and frequently assemble into short contigs, leaving many true within-genome links unobserved and causing viral bins to fragment. RESULTS: We present VirBinn, a graph-diffusion framework for viral binning from metagenomic Hi-C. VirBinn enhances virus-virus connectivity through two complementary mechanisms: random-walk-with-restart enhancement on the sparse virus-virus contact graph and host-guided diffusion that propagates viral seeds through the host network to infer indirect virus-virus associations. The enhanced views are integrated and clustered using Leiden community detection to produce viral metagenome-assembled genomes (vMAGs). On dataset-specific simulation benchmarks with ground truth, VirBinn consistently recovers more high-quality vMAGs than Hi-C-based and shotgun-based baselines and substantially increases the number of near-complete genomes. On four real metagenomic Hi-C datasets spanning human gut, pig gut, sheep gut (long-read assembly), and wastewater, VirBinn yields more high-completeness vMAGs under CheckV and produces bins with strong within-cluster contact support. Finally, host linkage analysis using reconstructed host MAGs reveals habitat-specific host-association patterns and plausible host taxonomic profiles. AVAILABILITY AND IMPLEMENTATION: VirBinn is available at https://github.com/dyxstat/VirBinn. The scripts to reproduce the results and figures in this article are available at https://github.com/dyxstat/Reproduce_VirBinn.

Genome, Viral↗

DEDB: a database of Drosophila melanogaster exons in splicing graph form.

BACKGROUND: A wealth of quality genomic and mRNA/EST sequences in recent years has provided the data required for large-scale genome-wide analysis of alternative splicing. We have capitalized on this by constructing a database that contains alternative splicing information organized as splicing graphs, where all transcripts arising from a single gene are collected, organized and classified. The splicing graph then serves as the basis for the classification of the various types of alternative splicing events. DESCRIPTION: DEDB http://proline.bic.nus.edu.sg/dedb/index.html is a database of Drosophila melanogaster exons obtained from FlyBase arranged in a splicing graph form that permits the creation of simple rules allowing for the classification of alternative splicing events. Pfam domains were also mapped onto the protein sequences allowing users to access the impact of alternative splicing events on domain organization. CONCLUSIONS: DEDB's catalogue of splicing graphs facilitates genome-wide classification of alternative splicing events for genome analysis. The splicing graph viewer brings together genome, transcript, protein and domain information to facilitate biologists in understanding the implications of alternative splicing.

Alternative Splicing↗

Transcriptome and genome conservation of alternative splicing events in humans and mice.

Combining mRNA and EST data in splicing graphs with whole genome alignments, we discover alternative splicing events that are conserved in both human and mouse transcriptomes. 1,964 of 19,156 (10%) loci examined contain one or more such alternative splicing events, with 2,698 total events. These events represent a lower bound on the amount of alternative splicing in the human genome. Also, as these alternative splicing events are conserved between the human and mouse transcriptomes they should be enriched for functionally significant alternative splicing events, free from much of the noise found in the EST libraries. Further classification of these alternative splicing events reveals that 1,037 (38.4%) are due to exon skipping, 497 (18.4%) are due to alternative 3' splice sites, 214 (7.9%) are due to alternative 5' splice sites, 75 (2.8%) are due to intron retention and the other 875 (32.4%) are due to other, more complicated, alternative splicing events. In addition, genomic sequences nearby these alternative splicing events display increased sequence conservation. Both the alternatively spliced exons and the proximal intron show increased levels of genomic conservation relative to constitutively spliced exons. For exon skipping events both intron regions flanking the exon are conserved while for alternative 5' and 3' splicing events the conservation is greater near the alternative splice site.

Algorithms↗

PangyPlot: multi-scale interactive visualization of pangenome variation graphs.

SUMMARY: Pangenome variation graphs integrate multiple samples into a unified representation, mitigating the reference bias inherent to linear genomes. However, these graphs can be large and structurally complex. Existing visualization tools are each confined to a fixed scale of resolution, requiring researchers to switch between multiple tools to examine variation at different levels of detail. PangyPlot is an interactive pangenome browser designed for multi-scale exploration of reference variation graphs from full chromosome to nucleotide-level sequence segments. PangyPlot anchors navigation to linear reference coordinates, organizes variation into hierarchical bubble structures, and uses a force-directed layout engine for automatic node arrangement. AVAILABILITY AND IMPLEMENTATION: An instance preloaded with data is available at https://pangyplot.research.sickkids.ca. Source code and documentation are openly available at https://github.com/strug-hub/pangyplot under the MIT License.

Software↗

Data note system for capturing laboratory data.

The complexity of genome data limits the usefulness of traditional database management systems. The highly interconnected structure of genome data can be captured in a data representation language based on the mathematical formalism of graphs. We have tailored graphs for describing genome data and have developed a database management system, called the Data Note System, for developing small databases to capture data from genome laboratories. To simplify the use of the Data Note System, a series of tools with graphical user interfaces has been developed. The system is designed to be easy to install and use by novice database developers with a minimal amount of computer expertise. We describe the tools and present examples of their use. The system consists of a storage facility, a schema editing tool to simplify the design of small databases, and three tools for data entry and querying.

Computers↗

gaftools: a toolkit for analyzing and manipulating pangenome alignments.

MOTIVATION: Linear reference genomes are ubiquitously used in genomics research, despite known biases associated with their use. In recent years, there has been a shift towards graph-based reference genomes to address some of these biases, which has required development of new algorithms and file formats. This has created a necessity for new tools capable of utilizing these formats and performing operations similar to those carried out by traditional methods. RESULTS: In this paper we present "gaftools," a multi-purpose tool that introduces several utilities for processing graph alignments in GAF format. gaftools enables users to index and sort alignments, with graph ordering serving as a necessary step for the sorting process. Additionally, it allows users to view subsets of alignments and perform realignment using the wavefront alignment algorithm, among other features. Many of these functionalities are inspired by SAMtools, which provides similar operations for linear genomes, while gaftools adapts and extends them for pangenomes. AVAILABILITY: gaftools is available under MIT license at https://github.com/marschall-lab/gaftools.

Software↗

SPC: a SPectral Component approach leveraging Identity-by-Descent graphs to address recent population structure in genomic analysis.

Population structure is a well-known confounder in statistical genetics, particularly in genome-wide association studies (GWAS), where it can lead to inflated test statistics and spurious associations. Traditional methods, such as principal components (PCs), commonly used to adjust for population structure, are limited in capturing fine-scale, non-linear patterns that arise from recent demographic events - patterns that are crucial for understanding rare variant effects. To address this challenge, we propose a novel method called SPectral Components (SPCs), which leverages identity-by-descent (IBD) graphs to capture and transform local, non-linear fine-scale population structure into continuous representations that can be seamlessly integrated into genetic analysis pipelines. Using both simulated datasets and empirical data from the UK Biobank (N ≈ 420,000), we demonstrate that SPCs outperform PCs in adjusting for fine-scale population structure. In simulations, SPCs explained over 90% of the fine-scale population structure with fewer components, while PCs captured less than 5%. In the UK Biobank, SPCs reduced the inflation of p-values in the GWAS of an environmental-driven phenotype by 12% compared to PCs, while maintaining a similar performance to PCs in height, a highly heritable phenotype. Additionally, SPCs improved rare variant association analyses, reducing genomic inflation (e.g., from 7.6 to 1.2 in one analysis), and provided more accurate heritability estimates. Spatial autocorrelation analysis further confirmed the ability of SPCs to account for environmental effects, reducing Moran's I for both environmental and heritable phenotypes more effectively than PCs. Overall, our findings demonstrate that SPCs provide a robust, scalable adjustment for recent population structure, offering a powerful alternative or complement to PCs in large-scale biobank studies.

GWAS↗

Pairwise graph edit distance characterizes the impact of the construction method on pangenome graphs.

MOTIVATION: Pangenome variation graphs are an increasingly used tool to perform genome analysis, aiming to replace a linear reference in a wide variety of genomic analyses. The construction of a variation graph from a collection of chromosome-size genome sequences is a difficult task that is generally addressed using a number of heuristics. The question that arises is to what extent the construction method influences the resulting graph, and the characterization of variability. RESULTS: We aim to characterize the differences between variation graphs derived from the same set of genomes with a metric which expresses and pinpoint differences. We designed a pairwise variation graph comparison algorithm, which establishes an edit distance between variation graphs, threading the genomes through both graphs. We applied our method to pangenome graphs built from yeast and human chromosome collections, and demonstrate that our method effectively characterizes discordances between pangenome graph construction methods and scales to real datasets. AVAILABILITY AND IMPLEMENTATION: pancat compare is published as free Rust software under the AGPL3.0 open source license. Source code and documentation are available at https://github.com/dubssieg/rs-pancat-compare. Snapshot available on Software Heritage at swh:1:dir:61acda8ba3dac1709ed60530147d3871831be629.

Algorithms↗

Syntons, metabolons and interactons: an exact graph-theoretical approach for exploring neighbourhood between genomic and functional data.

MOTIVATION: Modern comparative genomics does not restrict to sequence but involves the comparison of metabolic pathways or protein-protein interactions as well. Central in this approach is the concept of neighbourhood between entities (genes, proteins, chemical compounds). Therefore there is a growing need for new methods aiming at merging the connectivity information from different biological sources in order to infer functional coupling. RESULTS: We present a generic approach to merge the information from two or more graphs representing biological data. The method is based on two concepts. The first one, the correspondence multigraph, precisely defines how correspondence is performed between the primary data-graphs. The second one, the common connected components, defines which property of the multigraph is searched for. Although this problem has already been informally stated in the past few years, we give here a formal and general statement together with an exact algorithm to solve it. AVAILABILITY: The algorithm presented in this paper has been implemented in C. Source code is freely available for download at: http://www.inrialpes.fr/helix/people/viari/cccpart.

Algorithms↗

Genome rearrangements in mammalian evolution: lessons from human and mouse genomes.

Although analysis of genome rearrangements was pioneered by Dobzhansky and Sturtevant 65 years ago, we still know very little about the rearrangement events that produced the existing varieties of genomic architectures. The genomic sequences of human and mouse provide evidence for a larger number of rearrangements than previously thought and shed some light on previously unknown features of mammalian evolution. In particular, they reveal that a large number of microrearrangements is required to explain the differences in draft human and mouse sequences. Here we describe a new algorithm for constructing synteny blocks, study arrangements of synteny blocks in human and mouse, derive a most parsimonious human-mouse rearrangement scenario, and provide evidence that intrachromosomal rearrangements are more frequent than interchromosomal rearrangements. Our analysis is based on the human-mouse breakpoint graph, which reveals related breakpoints and allows one to find a most parsimonious scenario. Because these graphs provide important insights into rearrangement scenarios, we introduce a new visualization tool that allows one to view breakpoint graphs superimposed with genomic dot-plots.

Algorithms↗

Statistical Viewer: a tool to upload and integrate linkage and association data as plots displayed within the Ensembl genome browser.

BACKGROUND: To facilitate efficient selection and the prioritization of candidate complex disease susceptibility genes for association analysis, increasingly comprehensive annotation tools are essential to integrate, visualize and analyze vast quantities of disparate data generated by genomic screens, public human genome sequence annotation and ancillary biological databases. We have developed a plug-in package for Ensembl called "Statistical Viewer" that facilitates the analysis of genomic features and annotation in the regions of interest defined by linkage analysis. RESULTS: Statistical Viewer is an add-on package to the open-source Ensembl Genome Browser and Annotation System that displays disease study-specific linkage and/or association data as 2 dimensional plots in new panels in the context of Ensembl's Contig View and Cyto View pages. An enhanced upload server facilitates the upload of statistical data, as well as additional feature annotation to be displayed in DAS tracts, in the form of Excel Files. The Statistical View panel, drawn directly under the ideogram, illustrates lod score values for markers from a study of interest that are plotted against their position in base pairs. A module called "Get Map" easily converts the genetic locations of markers to genomic coordinates. The graph is placed under the corresponding ideogram features a synchronized vertical sliding selection box that is seamlessly integrated into Ensembl's Contig- and Cyto- View pages to choose the region to be displayed in Ensembl's "Overview" and "Detailed View" panels. To resolve Association and Fine mapping data plots, a "Detailed Statistic View" plot corresponding to the "Detailed View" may be displayed underneath. CONCLUSION: Features mapping to regions of linkage are accentuated when Statistic View is used in conjunction with the Distributed Annotation System (DAS) to display supplemental laboratory information such as differentially expressed disease genes in private data tracks. Statistic View is a novel and powerful visual feature that enhances Ensembl's utility as valuable resource for integrative genomic-based approaches to the identification of candidate disease susceptibility genes. At present there are no other tools that provide for the visualization of 2-dimensional plots of quantitative data scores against genomic coordinates in the context of a primary public genome annotation browser.

Chromosome Mapping↗

Reference-Free Variant Calling with Local Graph Construction with ska lo (SKA).

The study of genomic variants is increasingly important for public health surveillance of pathogens. Traditional variant-calling methods from whole-genome sequencing data rely on reference-based alignment, which can introduce biases and require significant computational resources. Alignment- and reference-free approaches offer an alternative by leveraging k-mer-based methods, but existing implementations often suffer from sensitivity limitations, particularly in high mutation density genomic regions. Here, we present ska lo, a graph-based algorithm that aims to identify within-strain variants in pathogen whole-genome sequencing data by traversing a colored De Bruijn graph and building variant groups (i.e. sets of variant combinations). Through in silico benchmarking and real-world dataset analyses, we demonstrate that ska lo achieves high sensitivity in single-nucleotide polymorphism (SNP) calls while also enabling the detection of insertions and deletions, as well as SNP positioning on a reference genome for recombination analyses. These findings highlight ska lo as a simple, fast, and effective tool for pathogen genomic epidemiology, extending the range of reference-free variant-calling approaches. ska lo is freely available as part of the SKA program (https://github.com/bacpop/ska.rust).

Polymorphism, Single Nucleotide↗

A versatile image analysis approach for simultaneous chromosome identification and localization of FISH probes.

Modern cytogenetic techniques, such as comparative genomic hybridization (CGH) and the multi-color fluorescence in situ hybridization (FISH) techniques of multiplex fluorescence in situ hybridization (M-FISH) and spectral karyotyping (SKY), require a coordinated banding analysis to maximize their usefulness. All of the methods currently used, including Giemsa (G-) banding, Alu banding, and 4',6-diamidino-2-phenyl-indole (DAPI) banding, have serious drawbacks. A simple and effective method to band chromosomes concurrently with FISH is needed. To address this problem, we stained chromosomes with DAPI and chromomycin A3, and then used an image analysis program to generate banding by dividing the image taken with a DAPI excitation filter by the image taken with a chromomycin A3 excitation filter. The result was a metaphase spread in which the chromosomes possessed a banding pattern characteristic of R-banding. The image analysis program was then used to generate linescans of pixel intensity versus relative position along the length of chromosomes that were banded using this technique, which we have called D/C R-banding. Each chromosome in a genome was represented by a characteristic scan profile, which was unaffected by FISH signals. Reference linescans were prepared by karyotyping D/C R-banded chromosomes for a given species, and then drawing lines along the length of the known chromosomes. The linescans were combined into a spreadsheet database, which was linked by dynamic data exchange to the image analysis program and normalized for length and intensity. The linescan of an unknown chromosome was then transferred to the spreadsheet, where it was normalized for length and intensity and overlaid on the linescans of each chromosome in the genome. Unknown chromosomes were identified by comparison of their graphs with graphs in the standardized reference genome. We have used this approach to create reference linescan karyotypes of several species, and to identify chromosomes on which FISH was performed.

Animals↗

Viewing genome data as objects for application development.

Genomics is becoming a data-intensive science, and an increasing number of laboratories are generating data which swamps storage in traditional paper-and-ink notebooks. Capturing the data flow requires large systems with multiple applications manipulating the same or similar data. Large systems often have conflicting requirements for data representation. Consistency across applications is a prime consideration, and appropriate data representation is an important issue in developing practical systems for molecular biologists. Graphs are a natural representation for describing genome data, while objects are good for modeling the behavior necessary for laboratory applications. We present a method for translating graph descriptions of genome data into objects using objects as views on graphs. Graph representations describe genome concepts while objects capture individual views for application development insuring consistency across genome applications.

Chromosome Mapping↗

Phylogenetic distribution and longitudinal persistence of plasmids in Mycobacterium abscessus.

Mycobacterium abscessus, a non-tuberculous mycobacterium, is a cause of severe respiratory infections, notably in individuals with underlying lung conditions. Its high levels of intrinsic and acquired antimicrobial resistance make it particularly difficult to treat and horizontally acquired genetic elements may facilitate the spread of resistance. A small number of plasmids have been identified in this species, but their distribution, transmission dynamics across subspecies and clonal lineages remain poorly characterized. We analysed short-read genomic data from 3,060 M. abscessus isolates, including longitudinal samples, to characterize plasmid diversity and dynamics. Using a graph-based pan-genome approach, we identified 28 plasmids, including 15 previously unreported plasmids, mapped their distribution onto the species phylogeny and assessed their functional potential. Overall, 23.1% of isolates carried at least one plasmid, with higher prevalence in dominant circulating clones (DCCs) compared with non-DCCs. Plasmid carriage varied across subspecies and clonal backgrounds, and plasmids encoded numerous genes which may be linked to bacterial adaptation. Several plasmids persisted across multiple time points within individual patients, suggesting they can be highly stable over the course of a chronic infection.

Plasmids↗

A heuristic graph comparison algorithm and its application to detect functionally related enzyme clusters.

The availability of computerized knowledge on biochemical pathways in the KEGG database opens new opportunities for developing computational methods to characterize and understand higher level functions of complete genomes. Our approach is based on the concept of graphs; for example, the genome is a graph with genes as nodes and the pathway is another graph with gene products as nodes. We have developed a simple method for graph comparison to identify local similarities, termed correlated clusters, between two graphs, which allows gaps and mismatches of nodes and edges and is especially suitable for detecting biological features. The method was applied to a comparison of the complete genomes of 10 microorganisms and the KEGG metabolic pathways, which revealed, not surprisingly, a tendency for formation of correlated clusters called FRECs (functionally related enzyme clusters). However, this tendency varied considerably depending on the organism. The relative number of enzymes in FRECs was close to 50% for Bacillus subtilis and Escherichia coli, but was <10% for SYNECHOCYSTIS: and Saccharomyces cerevisiae. The FRECs collection is reorganized into a collection of ortholog group tables in KEGG, which represents conserved pathway motifs with the information about gene clusters in all the completely sequenced genomes.

Algorithms↗

Identifying repeat domains in large genomes.

We present a graph-based method for the analysis of repeat families in a repeat library. We build a repeat domain graph that decomposes a repeat library into repeat domains, short subsequences shared by multiple repeat families, and reveals the mosaic structure of repeat families. Our method recovers documented mosaic repeat structures and suggests additional putative ones. Our method is useful for elucidating the evolutionary history of repeats and annotating de novo generated repeat libraries.

Algorithms↗

Identifying loci under positive selection in complex population histories.

Detailed modeling of a species' history is of prime importance for understanding how natural selection operates over time. Most methods designed to detect positive selection along sequenced genomes, however, use simplified representations of past histories as null models of genetic drift. Here, we present the first method that can detect signatures of strong local adaptation across the genome using arbitrarily complex admixture graphs, which are typically used to describe the history of past divergence and admixture events among any number of populations. The method-called graph-aware retrieval of selective sweeps (GRoSS)-has good power to detect loci in the genome with strong evidence for past selective sweeps and can also identify which branch of the graph was most affected by the sweep. As evidence of its utility, we apply the method to bovine, codfish, and human population genomic data containing panels of multiple populations related in complex ways. We find new candidate genes for important adaptive functions, including immunity and metabolism in understudied human populations, as well as muscle mass, milk production, and tameness in specific bovine breeds. We are also able to pinpoint the emergence of large regions of differentiation owing to inversions in the history of Atlantic codfish.

Animals↗