PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “genomic data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

BioViews: Java-based tools for genomic data visualization.

Visualization tools for bioinformatics ideally should provide universal access to the most current data in an interactive and intuitive graphical user interface. Since the introduction of Java, a language designed for distributed programming over the Web, the technology now exists to build a genomic data visualization tool that meets these requirements. Using Java we have developed a prototype genome browser applet (BioViews) that incorporates a three-level graphical view of genomic data: a physical map, an annotated sequence map, and a DNA sequence display. Annotated biological features are displayed on the physical and sequence-based maps, and the different views are interconnected. The applet is linked to several databases and can retrieve features and display hyperlinked textual data on selected features. In addition to browsing genomic data, different types of analyses can be performed interactively and the results of these analyses visualized alongside prior annotations. Our genome browser is built on top of extensible, reusable graphic components specifically designed for bioinformatics. Other groups can (and do) reuse this work in various ways. Genome centers can reuse large parts of the genome browser with minor modifications, bioinformatics groups working on sequence analysis can reuse components to build front ends for analysis programs, and biology laboratories can reuse components to publish results as dynamic Web documents.

Animals↗

Viewing genome data as objects for application development.

Genomics is becoming a data-intensive science, and an increasing number of laboratories are generating data which swamps storage in traditional paper-and-ink notebooks. Capturing the data flow requires large systems with multiple applications manipulating the same or similar data. Large systems often have conflicting requirements for data representation. Consistency across applications is a prime consideration, and appropriate data representation is an important issue in developing practical systems for molecular biologists. Graphs are a natural representation for describing genome data, while objects are good for modeling the behavior necessary for laboratory applications. We present a method for translating graph descriptions of genome data into objects using objects as views on graphs. Graph representations describe genome concepts while objects capture individual views for application development insuring consistency across genome applications.

Chromosome Mapping↗

An evaluation of the current state of genomic data privacy protection technology and a roadmap for the future.

The incorporation of genomic data into personal medical records poses many challenges to patient privacy. In response, various systems for preserving patient privacy in shared genomic data have been developed and deployed. Although these systems de-identify the data by removing explicit identifiers (e.g., name, address, or Social Security number) and incorporate sound security design principles, they suffer from a lack of formal modeling of inferences learnable from shared data. This report evaluates the extent to which current protection systems are capable of withstanding a range of re-identification methods, including genotype-phenotype inferences, location-visit patterns, family structures, and dictionary attacks. For a comparative re-identification analysis, the systems are mapped to a common formalism. Although there is variation in susceptibility, each system is deficient in its protection capacity. The author discovers patterns of protection failure and discusses several of the reasons why these systems are susceptible. The analyses and discussion within provide guideposts for the development of next-generation protection methods amenable to formal proofs.

Computer Security↗

Genomic data visualization on the Web.

UNLABELLED: Many types of genomic data can be represented in matrix format, with rows corresponding to genes and columns corresponding to gene features. The heat map is a popular technique for visualizing such data, plotting the data on a two-dimensional grid and using a color scale to represent the magnitude of each matrix entry. Prism is a Web-based software tool for generating annotated heat map visualizations of genome-wide data quickly. The tool provides a selection of genome-specific annotation catalogs as well as a catalog upload capability. The heat maps generated are clickable, allowing the user to drill down to examine specific matrix entries, and gene annotations are linked to relevant genomic databases. AVAILABILITY: http://noble.gs.washington.edu/prism

Computer Graphics↗

Prediction of higher order functional networks from genomic data.

Post-genomics may be defined in different ways depending on how one views the challenges after the discovery of the genome. A traditional view is to follow the concept of the central dogma in molecular biology, namely from genome to transcriptome to proteome. Projects are ongoing to analyse gene expression profiles both at the mRNA and protein levels, and to catalogue protein 3D structure families, which will no doubt help the understanding of the information in the genome. However, once complete, such experimentally determined catalogues of genes, RNAs and proteins only tell us about the building blocks of life. They do not tell us much about how life operates as a system, such as higher order functional behaviours of the cell or the organism. Thus, an alternative view of post-genomics is to go up from the molecular level to the cellular level and eventually to still higher levels, i.e., the biological systems. Bioinformatics provides basic concepts as well as practical methods to integrate this view with the traditional view and to analyse complex interactions among building blocks and with dynamic environments.

Computational Biology↗

AskBeacon-performing genomic data exchange and analytics with natural language.

MOTIVATION: Enabling clinicians and researchers to directly interact with global genomic data resources by removing technological barriers is vital for medical genomics. AskBeacon enables large language models (LLMs) to be applied to securely shared cohorts via the Global Alliance for Genomics and Health Beacon protocol. By simply "asking" Beacon, actionable insights can be gained, analyzed, and made publication-ready. RESULTS: In the Parkinson's Progression Markers Initiative (PPMI), we use natural language to ask whether the sex-differences observed in Parkinson's disease are due to X-linked or autosomal markers. AskBeacon returns a publication-ready visualization showing that for PPMI the autosomal marker occurred 1.4 times more often in males with Parkinson's disease than females, compared to no differences for the X-linked marker. We evaluate commercial and open-weight LLM models, as well as different architectures to identify the best strategy for translating research questions to Beacon queries. AskBeacon implements extensive safety guardrails to ensure that genomic data is not exposed to the LLM directly, and that generated code for data extraction, analysis and visualization process is sanitized and hallucination resistant, so data cannot be leaked or falsified. AVAILABILITY AND IMPLEMENTATION: AskBeacon is available at https://github.com/aehrc/AskBeacon.

Genomics↗

Inferring a tumor progression model for neuroblastoma from genomic data.

PURPOSE: The knowledge of the key genomic events that are causal to cancer development and progression not only is invaluable for our understanding of cancer biology but also may have a direct clinical impact. The task of deciphering a model of tumor progression by requiring that it explains (or at least does not contradict) known clinical and molecular evidence can be very demanding, particularly for cancers with complex patterns of clinical and molecular evidence. MATERIALS AND METHODS: We formalize the process of model inference and show how a progression model for neuroblastoma (NB) can be inferred from genomic data. The core idea of our method is to translate the model of clonal cancer evolution to mathematical testable rules of inheritance. Seventy-eight NB samples in stages 1, 4S, and 4 were analyzed with array-based comparative genomic hybridization. RESULTS: The pattern of recurrent genomic alterations in NB is strongly stage dependent and it is possible to identify traces of tumor progression in this type of data. CONCLUSION: A tumor progression model for neuroblastoma is inferred, which is in agreement with clinical evidence, explains part of the heterogeneity of the clinical behavior observed for NB, and is compatible with existing empirical models of NB progression.

Child↗

Protein network inference from multiple genomic data: a supervised approach.

MOTIVATION: An increasing number of observations support the hypothesis that most biological functions involve the interactions between many proteins, and that the complexity of living systems arises as a result of such interactions. In this context, the problem of inferring a global protein network for a given organism, using all available genomic data about the organism, is quickly becoming one of the main challenges in current computational biology. RESULTS: This paper presents a new method to infer protein networks from multiple types of genomic data. Based on a variant of kernel canonical correlation analysis, its originality is in the formalization of the protein network inference problem as a supervised learning problem, and in the integration of heterogeneous genomic data within this framework. We present promising results on the prediction of the protein network for the yeast Saccharomyces cerevisiae from four types of widely available data: gene expressions, protein interactions measured by yeast two-hybrid systems, protein localizations in the cell and protein phylogenetic profiles. The method is shown to outperform other unsupervised protein network inference methods. We finally conduct a comprehensive prediction of the protein network for all proteins of the yeast, which enables us to propose protein candidates for missing enzymes in a biosynthesis pathway. AVAILABILITY: Softwares are available upon request.

Artificial Intelligence↗

Optimized multilayer perceptrons for molecular classification and diagnosis using genomic data.

MOTIVATION: Multilayer perceptrons (MLP) represent one of the widely used and effective machine learning methods currently applied to diagnostic classification based on high-dimensional genomic data. Since the dimensionalities of the existing genomic data often exceed the available sample sizes by orders of magnitude, the MLP performance may degrade owing to the curse of dimensionality and over-fitting, and may not provide acceptable prediction accuracy. RESULTS: Based on Fisher linear discriminant analysis, we designed and implemented an MLP optimization scheme for a two-layer MLP that effectively optimizes the initialization of MLP parameters and MLP architecture. The optimized MLP consistently demonstrated its ability in easing the curse of dimensionality in large microarray datasets. In comparison with a conventional MLP using random initialization, we obtained significant improvements in major performance measures including Bayes classification accuracy, convergence properties and area under the receiver operating characteristic curve (A(z)). SUPPLEMENTARY INFORMATION: The Supplementary information is available on http://www.cbil.ece.vt.edu/publications.htm

Biomarkers, Tumor↗

Estimating the tempo and mode of gene family evolution from comparative genomic data.

Comparison of whole genomes has revealed that changes in the size of gene families among organisms is quite common. However, there are as yet no models of gene family evolution that make it possible to estimate ancestral states or to infer upon which lineages gene families have contracted or expanded. In addition, large differences in family size have generally been attributed to the effects of natural selection, without a strong statistical basis for these conclusions. Here we use a model of stochastic birth and death for gene family evolution and show that it can be efficiently applied to multispecies genome comparisons. This model takes into account the lengths of branches on phylogenetic trees, as well as duplication and deletion rates, and hence provides expectations for divergence in gene family size among lineages. The model offers both the opportunity to identify large-scale patterns in genome evolution and the ability to make stronger inferences regarding the role of natural selection in gene family expansion or contraction. We apply our method to data from the genomes of five yeast species to show its applicability.

Evolution, Molecular↗

The GDB human genome data base anno 1993.

Version 5.0 of the Genome Data Base (GDB) was released in March 1993. This document describes some of the significant changes to the types of data which are stored within the GDB. In addition to handling a wider scope of data, the GDB 5.0 application software now supports the X-Windows protocol. Although the GDB software still remains the most widely utilized method for accessing the data, alternate methods of access are now available, including direct SQL (Structured Query Language) queries, FTP (Internet File Transfer Protocol), WAIS (Wide Area Information Server), and other tools produced by third-party developers.

Chromosome Mapping↗

How to get the most from fission yeast genome data: a report from the 2006 European Fission Yeast Meeting computing workshop.

A fission yeast computing workshop 'How to get the most from the fission yeast genome data' was run as a satellite to the European Fission Yeast Meeting. The broad aims of the workshop were to provide fission yeast bench biologists with a set of tools and protocols to query the fission yeast genome data in specific ways, in order to extract biologically meaningful information of interest, which can be tailored to the needs of individual research projects. A description of the workshop content is provided and a selection of the tools presented are reviewed.

Computational Biology↗

The genome data base (GDB)--a human gene mapping repository.

The types of gene mapping data and its organization in the Genome Data Base (GDB) recently established at Johns Hopkins Medical School are described. The database provides a continuous online environment for data perusal and editing and is used as the informatics core for running the human gene mapping workshops. Current development is primarily concentrated on extending the types of map object, means of defining map location, map storage and representation. Experimental data structures have been created that permit storage of any type of map information, physical or genetic.

Baltimore↗

Finding association rules on heterogeneous genome data.

A novel approach for discovery of knowledge from genome data, which has been recently watched with interest in the research area of database, is applied to finding unified rules spreading over sequence, structure, and function of protein. As the result of experiments using data extracted from PDB, SWISS-PROT, and PROSITE, some association rules stating sequential/structural/functional aspects of two kinds of endopeptidases were found.

Computer Simulation↗

Visualizing the genome: techniques for presenting human genome data and annotations.

BACKGROUND: In order to take full advantage of the newly available public human genome sequence data and associated annotations, biologists require visualization tools ("genome browsers") that can accommodate the high frequency of alternative splicing in human genes and other complexities. RESULTS: In this article, we describe visualization techniques for presenting human genomic sequence data and annotations in an interactive, graphical format. These techniques include: one-dimensional, semantic zooming to show sequence data alongside gene structures; color-coding exons to indicate frame of translation; adjustable, moveable tiers to permit easier inspection of a genomic scene; and display of protein annotations alongside gene structures to show how alternative splicing impacts protein structure and function. These techniques are illustrated using examples from two genome browser applications: the Neomorphic GeneViewer annotation tool and ProtAnnot, a prototype viewer which shows protein annotations in the context of genomic sequence. CONCLUSION: By presenting techniques for visualizing genomic data, we hope to provide interested software developers with a guide to what features are most likely to meet the needs of biologists as they seek to make sense of the rapidly expanding body of public genomic data and annotations.

Alternative Splicing↗

Mitochondrial genomics of gadine fishes: implications for taxonomy and biogeographic origins from whole-genome data sets.

Phylogenetic analysis of 13 substantially complete mitochondrial DNA genome sequences (14,036 bp) from 10 taxa of gadine codfishes and pollock provides highly corroborated resolution of outstanding questions on their biogeographic evolution. Of 6 resolvable nodes among species, 4 were supported by >95% of bootstrap replications in parsimony, distance, likelihood, and similarly high posterior probabilities in bayesian analyses, one by 85%-95% according to the method of analysis, and one by 99% by one method and a majority of the other two. The endemic Pacific species, walleye pollock (Theragra chalcogramma), is more closely related to the endemic Atlantic species, Atlantic cod (Gadus macrocephalus), than either is to a second Pacific endemic, Pacific cod (Gadus macrocephalus). The walleye pollock should thus be referred to the genus Gadus as originally described (Gadus chalcogrammus Pallas 1811). Arcto-Atlantic Greenland cod, previously regarded as a distinct species (G. ogac), are a genomically distinguishable subspecies within pan-Pacific G. macrocephalus. Of the 2 endemic Arctic Ocean genera, Polar cod (Boreogadus) as the outgroup to Arctic cod (Arctogadus) and Gadus sensu lato is more strongly supported than a pairing of Boreogadus and Arctogadus as sister taxa. Taking into consideration historical patterns of hydrogeography, we outline a hypothesis of the origin of the 2 endemic Pacific species as independent but simultaneous invasions through the Bering Strait from an Arcto-Atlantic ancestral lineage. In contrast to the genome data, the complete proteome sequence (3830 amino acids) resolved only 3 nodes with >95% confidence, and placed Alaska pollock outside the Gadus clade owing to reversal mutations in the ND5 locus that restore ancestral, non-Gadus, amino acid residues in that species.

Amino Acid Sequence↗

Utilizing logical relationships in genomic data to decipher cellular processes.

The wealth of available genomic data has spawned a corresponding interest in computational methods that can impart biological meaning and context to these experiments. Traditional computational methods have drawn relationships between pairs of proteins or genes based on notions of equality or similarity between their patterns of occurrence or behavior. For example, two genes displaying similar variation in expression, over a number of experiments, may be predicted to be functionally related. We have introduced a natural extension of these approaches, instead identifying logical relationships involving triplets of proteins. Triplets provide for various discrete kinds of logic relationships, leading to detailed inferences about biological associations. For instance, a protein C might be encoded within an organism if, and only if, two other proteins A and B are also both encoded within the organism, thus suggesting that gene C is functionally related to genes A and B. The method has been applied fruitfully to both phylogenetic and microarray expression data, and has been used to associate logical combinations of protein activity with disease state phenotypes, revealing previously unknown ternary relationships among proteins, and illustrating the inherent complexities that arise in biological data.

Algorithms↗

Scalable assembly of Ascaris mitogenomes from whole-genome data reveals a novel clade.

The genus Ascaris is an important group of giant parasitic roundworms, infecting over 700 million people globally and causing substantial economic losses in domestic pigs. Whilst species of Ascaris are morphologically indistinguishable, analysis of mitochondrial loci has revealed three clades (A, B, C) broadly associated with host species and geographic distribution. The diversity within these lineages may expand with the addition of further genomic data. Here, we present a bioinformatic framework for de novo assembly of complete mitochondrial genomes (mitogenomes) from low-coverage whole-genome data through host-read depletion or mtDNA read enrichment, followed by mtDNA-specific assembly. Our approach yielded 149 high-quality Ascaris mitogenome assemblies, enabling the study of population-level diversity, including the identification of a novel clade (Clade D, designated here) associated with human samples from Ethiopia. Our analysis further revealed Clade C to comprise of pig-derived samples from Europe based on characterisation of worms isolated in Germany. The methods described here provide a scalable framework for mitogenome reconstruction with insights into roundworm population-genomic and phylogenetic studies.

Animals↗