PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Detailed mapping of RNA secondary structures in core and NS5B-encoding region sequences of hepatitis C virus by RNase cleavage and novel bioinformatic prediction methods.

There is accumulating evidence from bioinformatic studies that hepatitis C virus (HCV) possesses extensive RNA secondary structure in the core and NS5B-encoding regions of the genome. Recent functional studies have defined one such stem-loop structure in the NS5B region as an essential cis-acting replication element (CRE). A program was developed (STRUCTUR_DIST) that analyses multiple rna-folding patterns predicted by mfold to determine the evolutionary conservation of predicted stem-loop structures and, by a new method, to analyse frequencies of covariant sites in predicted RNA folding between HCV genotypes. These novel bioinformatic methods have been combined with enzymic mapping of RNA transcripts from the core and NS5B regions to precisely delineate the RNA structures that are present in these genomic regions. Together, these methods predict the existence of multiple, often juxtaposed stem-loops that are found in all HCV genotypes throughout both regions, as well as several strikingly conserved single-stranded regions, one of which coincides with a region of the genome to which ribosomal access is required for translation initiation. Despite the existence of marked sequence conservation between genotypes in the HCV CRE and single-stranded regions, there was no evidence for comparable suppression of variability at either synonymous or non-synonymous sites in the other predicted stem-loop structures. The configuration and genetic variability of many of these other NS5B and core structures is perhaps more consistent with their involvement in genome-scale ordered RNA structure, a structural configuration of the genomes of many positive-stranded RNA viruses that is associated with host persistence.

Computational Biology↗

A bioinformatics-based strategy identifies c-Myc and Cdc25A as candidates for the Apmt mammary tumor latency modifiers.

The epistatically interacting modifier loci (Apmt1 and Apmt2) accelerate the polyoma Middle-T (PyVT)-induced mammary tumor. To identify potential candidate genes loci, a combined bioinformatics and genomics strategy was used. On the basis of the assumption that the loci were functioning in the same or intersecting pathways, a search of the literature databases was performed to identify molecular pathways containing genes from both candidate intervals. Among the genes identified by this method were the cell cycle-associated genes Cdc25A and c-Myc, both of which have been implicated in breast cancer. Genomic sequencing revealed noncoding polymorphism in both genes, in the promoter region of Cdc25A, and in the 3' UTR of c-Myc. Molecular and in vitro analysis showed that the polymorphisms were functionally significant. In vivo analysis was performed by generating compound PyVT/Myc double-transgenic animals to mimic the hypothetical model, and was found to recapitulate the age-of-onset phenotype. These data suggest that c-Myc and Cdc25A are Apmt1 and Apmt2, and suggest that, at least in certain instances, bioinformatics can be utilized to bypass congenic construction and subsequent mapping in conventional QTL studies.

3T3 Cells↗

BioMOBY successfully integrates distributed heterogeneous bioinformatics Web Services. The PlaNet exemplar case.

The burden of non-interoperability between on-line genomic resources is increasingly the rate-limiting step in large-scale genomic analysis. BioMOBY is a biological Web Service interoperability initiative that began as a retreat of representatives from the model organism database community in September, 2001. Its long-term goal is to provide a simple, extensible platform through which the myriad of on-line biological databases and analytical tools can offer their information and analytical services in a fully automated and interoperable way. Of the two branches of the larger BioMOBY project, the Web Services branch (MOBY-S) has now been deployed over several dozen data sources worldwide, revealing some significant observations about the nature of the integrative biology problem; in particular, that Web Service interoperability in the domain of bioinformatics is, unexpectedly, largely a syntactic rather than a semantic problem. That is to say, interoperability between bioinformatics Web Services can be largely achieved simply by specifying the data structures being passed between the services (syntax) even without rich specification of what those data structures mean (semantics). Thus, one barrier of the integrative problem has been overcome with a surprisingly simple solution. Here, we present a non-technical overview of the critical components that give rise to the interoperable behaviors seen in MOBY-S and discuss an exemplar case, the PlaNet consortium, where MOBY-S has been deployed to integrate the on-line plant genome databases and analytical services provided by a European consortium of databases and data service providers.

Computational Biology↗

Bioinformatic insights from metagenomics through visualization.

Cutting-edge biological and bioinformatics research seeks a systems perspective through the analysis of multiple types of high-throughput and other experimental data for the same sample. Systems-level analysis requires the integration and fusion of such data, typically through advanced statistics and mathematics. Visualization is a complementary computational approach that supports integration and analysis of complex data or its derivatives. We present a bioinformatics visualization prototype, Juxter, which depicts categorical information derived from or assigned to these diverse data for the purpose of comparing patterns across categorizations. The visualization allows users to easily discern correlated and anomalous patterns in the data. These patterns, which might not be detected automatically by algorithms, may reveal valuable information leading to insight and discovery. We describe the visualization and interaction capabilities and demonstrate its utility in a new field, metagenomics, which combines molecular biology and genetics to identify and characterize genetic material from multi-species microbial samples.

Algorithms↗

An insight-based methodology for evaluating bioinformatics visualizations.

High-throughput experiments, such as gene expression microarrays in the life sciences, result in very large data sets. In response, a wide variety of visualization tools have been created to facilitate data analysis. A primary purpose of these tools is to provide biologically relevant insight into the data. Typically, visualizations are evaluated in controlled studies that measure user performance on predetermined tasks or using heuristics and expert reviews. To evaluate and rank bioinformatics visualizations based on real-world data analysis scenarios, we developed a more relevant evaluation method that focuses on data insight. This paper presents several characteristics of insight that enabled us to recognize and quantify it in open-ended user tests. Using these characteristics, we evaluated five microarray visualization tools on the amount and types of insight they provide and the time it takes to acquire it. The results of the study guide biologists in selecting a visualization tool based on the type of their microarray data, visualization designers on the key role of user interaction techniques, and evaluators on a new approach for evaluating the effectiveness of visualizations for providing insight. Though we used the method to analyze bioinformatics visualizations, it can be applied to other domains.

Algorithms↗

Bioinformatic analysis of Helicobacter pylori XGPRTase: a potential therapeutic target.

BACKGROUND: Xanthine-guanine phosphoribosyltransferase (XGPRTase) is an enzyme of purine nucleotide salvage synthesis. The gpt gene of Helicobacter pylori has been annotated as encoding an XGPRTase and proposed as essential for survival of the bacterium in vitro. The aims of this work were to investigate the structure of H. pylori XGPRTase and to compare the key features of the enzyme to other phosphoribosyltransferases employing computational, modelling, and bioinformatic tools. MATERIALS AND METHODS: XGPRTase activity was measured in the cytosolic fraction of H. pylori by (31)P-nuclear magnetic resonance spectroscopy, and also in recombinant XGPRTase produced by a cell-free expression system. Bioinformatics was employed to analyze the phylogeny of XGPRTase, and a structural model of the XGPRTase was built using threading techniques. The observed interactions of purine phosphoribosyltransferases with immucillin-GP were used to study the theoretical interactions of H. pylori XGPRTase with this transition-state analog. RESULTS: It was demonstrated that the gpt gene of H. pylori encodes a functional XGPRTase enzyme. Analyses of the XGPRTase sequence showed that the enzyme is significantly divergent from equivalent mammalian enzymes. Modelling served to identify specific features of the enzyme and key residues involved in catalysis. CONCLUSIONS: The H. pylori XGPRTase is structurally similar to other phosphoribosyltransferase enzymes, but there were significant differences between the hood domain of H. pylori XGPRTase and other purine salvage phosphoribosyltransferases. Significant differences were found between the interactions of the H. pylori and human enzymes with a purine phosphoribosyltransferase inhibitor.

Amino Acid Sequence↗

Bioinformatic identification of tandem repeat antigens of the Leishmania donovani complex.

With large amounts of parasite gene sequence available, additional bioinformatic tools to screen these sequences for identifying genes encoding antigens are needed. Proteins containing tandem repeat (TR) domains are often B-cell antigens, and antibody responses toward TR domains of the proteins are dominant in human infected with certain parasites. We hypothesized that antigens of serological significance could be identified with a search for TR domains. Here we show the result of bioinformatic screening of the gene sequence database of the parasitic protozoan Leishmania infantum. Of 8,191 genes scanned, 64 genes contained TR domains. Of the 64 genes, 22 encoded previously characterized antigens; the remaining 42 genes were previously uncharacterized. By using sera from Sudanese visceral leishmaniasis patients, we confirmed that the TR domains of LinJ11.0070, LinJ25.1100, LinJ27.0400, and LinJ29.0110, which were from the 42 uncharacterized proteins, are also antigenic. The results suggest the validity of this approach for identifying leishmanial antigens of serological significance.

Animals↗

Genome-based bioinformatic selection of chromosomal Bacillus anthracis putative vaccine candidates coupled with proteomic identification of surface-associated antigens.

Bacillus anthracis (Ames strain) chromosome-derived open reading frames (ORFs), predicted to code for surface exposed or virulence related proteins, were selected as B. anthracis-specific vaccine candidates by a multistep computational screen of the entire draft chromosome sequence (February 2001 version, 460 contigs, The Institute for Genomic Research, Rockville, Md.). The selection procedure combined preliminary annotation (sequence similarity searches and domain assignments), prediction of cellular localization, taxonomical and functional screen and additional filtering criteria (size, number of paralogs). The reductive strategy, combined with manual curation, resulted in selection of 240 candidate ORFs encoding proteins with putative known function, as well as 280 proteins of unknown function. Proteomic analysis of two-dimensional gels of a B. anthracis membrane fraction, verified the expression of some gene products. Matrix-assisted laser desorption ionization-time-of-flight mass spectrometry analyses allowed identification of 38 spots cross-reacting with sera from B. anthracis immunized animals. These spots were found to represent eight in vivo immunogens, comprising of EA1, Sap, and 6 proteins whose expression and immunogenicity was not reported before. Five of these 8 immunogens were preselected by the bioinformatic analysis (EA1, Sap, 2 novel SLH proteins and peroxiredoxin/AhpC), as vaccine candidates. This study demonstrates that a combination of the bioinformatic and proteomic strategies may be useful in promoting the development of next generation anthrax vaccine.

Adhesins, Bacterial↗

Determining human immunodeficiency virus coreceptor use in a clinical setting: degree of correlation between two phenotypic assays and a bioinformatic model.

Two recombinant phenotypic assays for human immunodeficiency virus (HIV) coreceptor usage and an HIV envelope genotypic predictor were employed on a set of clinically derived HIV type 1 (HIV-1) samples in order to evaluate the concordance between measures. Previously genotyped HIV-1 samples derived from antiretroviral-naïve individuals were tested for coreceptor usage using two independent phenotyping methods. Phenotypes were determined by validated recombinant assays that incorporate either an approximately 2,500-bp ("Trofile" assay) or an approximately 900-bp (TRT assay) fragment of the HIV envelope gp120. Population-based HIV envelope V3 loop sequences ( approximately 105 bp) were derived by automated sequence analysis. Genotypic coreceptor predictions were performed using a support vector machine model trained on a separate genotype-Trofile phenotype data set. HIV coreceptor usage was obtained from both phenotypic assays for 74 samples, with an overall 85.1% concordance. There was no evidence of a difference in sensitivity between the two phenotypic assays. A bioinformatic algorithm based on a support vector machine using HIV V3 genotype data was able to achieve 86.5% and 79.7% concordance with the Trofile and TRT assays, respectively, approaching the degree of agreement between the two phenotype assays. In most cases, the phenotype assays and the bioinformatic approach gave similar results. However, in cases where there were differences in the tropism results, it was not clear which of the assays was "correct." X4 (CXCR4-using) minority species in clinically derived samples likely complicate the interpretation of both phenotypic and genotypic assessments of HIV tropism.

Computational Biology↗

Genomic and bioinformatics analyses of HAdV-4vac and HAdV-7vac, two human adenovirus (HAdV) strains that constituted original prophylaxis against HAdV-related acute respiratory disease, a reemerging epidemic disease.

Vaccine strains of human adenovirus serotypes 4 and 7 (HAdV-4vac and HAdV-7vac) have been used successfully to prevent adenovirus-related acute respiratory disease outbreaks. The genomes of these two vaccine strains have been sequenced, annotated, and compared with their prototype equivalents with the goals of understanding their genomes for molecular diagnostics applications, vaccine redevelopment, and HAdV pathoepidemiology. These reference genomes are archived in GenBank as HAdV-4vac (35,994 bp; AY594254) and HAdV-7vac (35,240 bp; AY594256). Bioinformatics and comparative whole-genome analyses with their recently reported and archived prototype genomes reveal six mismatches and four insertions-deletions (indels) between the HAdV-4 prototype and vaccine strains, in contrast to the 611 mismatches and 130 indels between the HAdV-7 prototype and vaccine strains. Annotation reveals that the HAdV-4vac and HAdV-7vac genomes contain 51 and 50 coding units, respectively. Neither vaccine strain appears to be attenuated for virulence based on bioinformatics analyses. There is evidence of genome recombination, as the inverted terminal repeat of HAdV-4vac is initially identical to that of species C whereas the prototype is identical to species B1. These vaccine reference sequences yield unique genome signatures for molecular diagnostics. As a molecular forensics application, these references identify the circulating and problematic 1950s era field strains as the original HAdV-4 prototype and the Greider prototype, from which the vaccines are derived. Thus, they are useful for genomic comparisons to current epidemic and reemerging field strains, as well as leading to an understanding of pathoepidemiology among the human adenoviruses.

Acute Disease↗

Genomic and bioinformatics analysis of HAdV-4, a human adenovirus causing acute respiratory disease: implications for gene therapy and vaccine vector development.

Human adenovirus serotype 4 (HAdV-4) is a reemerging viral pathogenic agent implicated in epidemic outbreaks of acute respiratory disease (ARD). This report presents a genomic and bioinformatics analysis of the prototype 35,990-nucleotide genome (GenBank accession no. AY594253). Intriguingly, the genome analysis suggests a closer phylogenetic relationship with the chimpanzee adenoviruses (simian adenoviruses) rather than with other human adenoviruses, suggesting a recent origin of HAdV-4, and therefore species E, through a zoonotic event from chimpanzees to humans. Bioinformatics analysis also suggests a pre-zoonotic recombination event, as well, between species B-like and species C-like simian adenoviruses. These observations may have implications for the current interest in using chimpanzee adenoviruses in the development of vectors for human gene therapy and for DNA-based vaccines. Also, the reemergence, surveillance, and treatment of HAdV-4 as an ARD pathogen is an opportunity to demonstrate the use of genome determination as a tool for viral infectious disease characterization and epidemic outbreak surveillance: for example, rapid and accurate low-pass sequencing and analysis of the genome. In particular, this approach allows the rapid identification and development of unique probes for the differentiation of family, species, serotype, and strain (e.g., pathogen genome signatures) for monitoring epidemic outbreaks of ARD.

Adenovirus Infections, Human↗

The clinical bioinformatics ontology: a curated semantic network utilizing RefSeq information.

Existing medical vocabularies lack rich terms to describe findings that are generated by modem molecular diagnostic procedures. Most bioinformatics resources were designed primarily to support the needs of the research community. We describe the development of a curated resource, the Clinical Bioinformatics Ontology (CBO), a semantic network appropriate for describing clinically significant genomics concepts. The CBO includes concepts appropriate for both molecular diagnostics and cytogenetics. A standardized methodology based on consistent application of RefSeq information is applied to the curation of the CBO in order to provide a reproducible and reliable tool. Challenges related to this curation process are discussed in this paper. At the time of submission the CBO included 4,069 concepts, associated by 8,463 relationships.

Computational Biology↗

Fuzzy logic in medicine and bioinformatics.

The purpose of this paper is to present a general view of the current applications of fuzzy logic in medicine and bioinformatics. We particularly review the medical literature using fuzzy logic. We then recall the geometrical interpretation of fuzzy sets as points in a fuzzy hypercube and present two concrete illustrations in medicine (drug addictions) and in bioinformatics (comparison of genomes).

Journal Article↗

Practical bioinformatics for proteomics.

We used various bioinformatic tools to examine the unknown protein gi|12841975 that was up-regulated in mouse diabetic kidneys. The data indicate that this unknown protein is, indeed, the PEBP. Motif scanning showed that this protein contains several kinase motifs, especially PKC that plays an important role in the pathogenesis of diabetic nephropathy [24, 25]. We therefore hypothesize that this protein (PEBP) has a potential functional role in PKC-dependent pathogenic pathways of diabetic nephropathy. Further study will be focused on phosphorylation pathways of the PEBP and its substrates. In summary, we have presented a case study that outlines our approach to further characterize the unknown proteins identified by peptide mass fingerprinting. Publicly accessible bioinformatic tools can provide a wealth of information to guide subsequent approaches that use traditional molecular biology tools.

Animals↗

Overview of bioinformatics and its application to oral genomics.

The "informatics revolution" in both bioinformatics and dental informatics will eventually change the way we practice dentistry. This convergence will play a pivotal role in creating a bridge of opportunity by integrating scientific and clinical specialties to promote the advances in treatment, risk assessment, diagnosis, therapeutics, and oral health-care outcome. Bioinformatics has been an emerging field in the biomedical research community and has been gaining momentum in dental medicine. This area has created a steady stream of large and complex genomic data, which has transformed the way a clinical or basic science researcher approaches genomic research. This application to dental medicine, termed "oral genomics", can aid in the molecular understanding of the genes and proteins, their interactions, pathways, and networks that are responsible for the development and progression of oral diseases and disorders. As the result of the Human Genome Project, new advances have prompted high-throughput technologies, such as DNA microarrays, which have become accepted tools in the biomedical research community. This manuscript reviews the two most commonly used microarray technologies, basic microarray data analysis, and the results from several ongoing oral cancer genomic studies.

Carcinoma, Squamous Cell↗

SeqHound: biological sequence and structure database as a platform for bioinformatics research.

BACKGROUND: SeqHound has been developed as an integrated biological sequence, taxonomy, annotation and 3-D structure database system. It provides a high-performance server platform for bioinformatics research in a locally-hosted environment. RESULTS: SeqHound is based on the National Center for Biotechnology Information data model and programming tools. It offers daily updated contents of all Entrez sequence databases in addition to 3-D structural data and information about sequence redundancies, sequence neighbours, taxonomy, complete genomes, functional annotation including Gene Ontology terms and literature links to PubMed. SeqHound is accessible via a web server through a Perl, C or C++ remote API or an optimized local API. It provides functionality necessary to retrieve specialized subsets of sequences, structures and structural domains. Sequences may be retrieved in FASTA, GenBank, ASN.1 and XML formats. Structures are available in ASN.1, XML and PDB formats. Emphasis has been placed on complete genomes, taxonomy, domain and functional annotation as well as 3-D structural functionality in the API, while fielded text indexing functionality remains under development. SeqHound also offers a streamlined WWW interface for simple web-user queries. CONCLUSIONS: The system has proven useful in several published bioinformatics projects such as the BIND database and offers a cost-effective infrastructure for research. SeqHound will continue to develop and be provided as a service of the Blueprint Initiative at the Samuel Lunenfeld Research Institute. The source code and examples are available under the terms of the GNU public license at the Sourceforge site http://sourceforge.net/projects/slritools/ in the SLRI Toolkit.

Amino Acid Sequence↗

A web services choreography scenario for interoperating bioinformatics applications.

BACKGROUND: Very often genome-wide data analysis requires the interoperation of multiple databases and analytic tools. A large number of genome databases and bioinformatics applications are available through the web, but it is difficult to automate interoperation because: 1) the platforms on which the applications run are heterogeneous, 2) their web interface is not machine-friendly, 3) they use a non-standard format for data input and output, 4) they do not exploit standards to define application interface and message exchange, and 5) existing protocols for remote messaging are often not firewall-friendly. To overcome these issues, web services have emerged as a standard XML-based model for message exchange between heterogeneous applications. Web services engines have been developed to manage the configuration and execution of a web services workflow. RESULTS: To demonstrate the benefit of using web services over traditional web interfaces, we compare the two implementations of HAPI, a gene expression analysis utility developed by the University of California San Diego (UCSD) that allows visual characterization of groups or clusters of genes based on the biomedical literature. This utility takes a set of microarray spot IDs as input and outputs a hierarchy of MeSH Keywords that correlates to the input and is grouped by Medical Subject Heading (MeSH) category. While the HTML output is easy for humans to visualize, it is difficult for computer applications to interpret semantically. To facilitate the capability of machine processing, we have created a workflow of three web services that replicates the HAPI functionality. These web services use document-style messages, which means that messages are encoded in an XML-based format. We compared three approaches to the implementation of an XML-based workflow: a hard coded Java application, Collaxa BPEL Server and Taverna Workbench. The Java program functions as a web services engine and interoperates with these web services using a web services choreography language (BPEL4WS). CONCLUSION: While it is relatively straightforward to implement and publish web services, the use of web services choreography engines is still in its infancy. However, industry-wide support and push for web services standards is quickly increasing the chance of success in using web services to unify heterogeneous bioinformatics applications. Due to the immaturity of currently available web services engines, it is still most practical to implement a simple, ad-hoc XML-based workflow by hard coding the workflow as a Java application. For advanced web service users the Collaxa BPEL engine facilitates a configuration and management environment that can fully handle XML-based workflow.

Computational Biology↗

libcov: a C++ bioinformatic library to manipulate protein structures, sequence alignments and phylogeny.

BACKGROUND: An increasing number of bioinformatics methods are considering the phylogenetic relationships between biological sequences. Implementing new methodologies using the maximum likelihood phylogenetic framework can be a time consuming task. RESULTS: The bioinformatics library libcov is a collection of C++ classes that provides a high and low-level interface to maximum likelihood phylogenetics, sequence analysis and a data structure for structural biological methods. libcov can be used to compute likelihoods, search tree topologies, estimate site rates, cluster sequences, manipulate tree structures and compare phylogenies for a broad selection of applications. CONCLUSION: Using this library, it is possible to rapidly prototype applications that use the sophistication of phylogenetic likelihoods without getting involved in a major software engineering project. libcov is thus a potentially valuable building block to develop in-house methodologies in the field of protein phylogenetics.

Algorithms↗