PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

A new bioinformatic approach to detect common 3D sites in protein structures.

An innovative bioinformatic method has been designed and implemented to detect similar three-dimensional (3D) sites in proteins. This approach allows the comparison of protein structures or substructures and detects local spatial similarities: this method is completely independent from the amino acid sequence and from the backbone structure. In contrast to already existing tools, the basis for this method is a representation of the protein structure by a set of stereochemical groups that are defined independently from the notion of amino acid. An efficient heuristic for finding similarities that uses graphs of triangles of chemical groups to represent the protein structures has been developed. The implementation of this heuristic constitutes a software named SuMo (Surfing the Molecules), which allows the dynamic definition of chemical groups, the selection of sites in the proteins, and the management and screening of databases. To show the relevance of this approach, we focused on two extreme examples illustrating convergent and divergent evolution. In two unrelated serine proteases, SuMo detects one common site, which corresponds to the catalytic triad. In the legume lectins family composed of >100 structures that share similar sequences and folds but may have lost their ability to bind a carbohydrate molecule, SuMo discriminates between functional and non-functional lectins with a selectivity of 96%. The time needed for searching a given site in a protein structure is typically 0.1 s on a PIII 800MHz/Linux computer; thus, in further studies, SuMo will be used to screen the PDB.

Algorithms↗

Use of bioinformatics tools for the annotation of disease-associated mutations in animal models.

Single-point mutations are one of the most frequent causes of genetic variability in both human and close species. The recent availability of different bioinformatics tools for annotating human single nucleotide polymorphisms (SNPs) has opened the possibility of using them to score SNPs from species with a biomedical interest, in particular from mice and other models of human disease. Also, this ability to predict pathogenicity of single point mutations in one species, based on data from another species, opens the possibility to predict the pathological character of single point mutations in humans using data from well-characterized model systems of human disease. This could provide a valuable alternative to the more traditional genetic population approaches. However, transferral of prediction tools may be limited by different factors, from a species bias in the training set, to a large sequence divergence between the proteomes of the training and the target species. Here we study the conditions under which prediction tools can be transferred among species, concentrating in the case of mice. We find that for the majority of the human-mouse homolog pairs, the sequence similarity is large enough to preserve the pathological character of mutations among species, in general. We then establish that prediction/annotation tools developed for one organism can be used to predict the neutral/pathological character of mutations/SNPs in the other organism.

Animals↗

Post-genomic virology: the impact of bioinformatics, microarrays and proteomics on investigating host and pathogen interactions.

Post-genomic research encompasses many diverse aspects of modern science. These include the two broad subject areas of computational biology (bioinformatics) and functional genomics. Laboratory based functional genomics aims to measure and assess either the messenger RNA (mRNA) levels (transcriptome studies) or the protein content (proteome studies) of cells and tissues. All of these methods have been applied recently to the study of host and pathogen interactions for both bacteria and viruses. A basic overview of the technology is given in this review together with approaches to data analysis. The wealth of information produced from even these preliminary studies has shown the generalities, subtleties and specificities of host-pathogen interactions. Such research should ultimately result in new methods for diagnosing and treating infectious diseases.

Amino Acid Sequence↗

Identification of a novel human glutathione S-transferase using bioinformatics.

In searching the expressed sequence tag (EST) data-base of GenBank with coding sequences of 11 known human glutathione S-transferases in conjunction with bioinformatic analysis, we have identified five ESTs that encode a new human glutathione S-transferase (GST) designated GST A4. The cDNA clone (I.M.A.G.E. Consortium cDNA Clone ID 515157) had an insert length of 1279 bp and contains an open reading frame of 666 bp, which encodes a protein of 222 amino acid residues. The GST A4 protein is identical in length to human GST A1 and A2 and is 54% identical to human GST A1 and A2. Sequence comparison with other human GSTs suggests that it is a new GST belonging to the alpha class GSTs. Northern blot analysis and EST database searches have demonstrated that the GST A4 mRNA is expressed at a high level in brain, placenta, and skeletal muscle and much lower in lung and liver. Analysis of the sequence tagged site (STS) database indicated that the GST A4 gene is located on chromosome 6. This STS represents a previously unidentified transcript further confirming the novelty of the new sequence.

Amino Acid Sequence↗

Biomarkers of human skin cells identified using DermArray DNA arrays and new bioinformatics methods.

Biomarker genes of human skin-derived cells were identified by new simple bioinformatic methods and DNA microarray analysis utilizing in vitro cultures of normal neonatal human epidermal keratinocytes, melanocytes, and dermal fibroblasts. A survey of 4405 human cDNAs was performed using DermArray DNA microarrays. Biomarkers were rank ordered by "likelihood ratio" algorithms and stringent selection criteria that have general applicability for analyzing a minimum of three RNA samples. Signature biomarker genes (up-regulated in one cell type) and anti-signature biomarker genes (down-regulated in one cell type) were determined for the three major skin cell types. Many of the signature genes are known biomarkers for these cell types. In addition, 17 signature genes were identified as ESTs, and 22 anti-signature biomarkers were discovered. Quantitative RT-PCR was used to verify nine signature biomarker genes. A total of 158 biomarkers of normal human skin cells were identified, many of which may be valuable in diagnostic applications and as molecular targets for drug discovery and therapeutic intervention.

Algorithms↗

Molecular complementarity III. peptide complementarity as a basis for peptide receptor evolution: a bioinformatic case study of insulin, glucagon and gastrin.

Dwyer has suggested that peptide receptors evolved from self-aggregating peptides so that peptide receptors should incorporate regions of high homology with the peptide ligand. If one considers self-aggregation to be a particular manifestation of molecular complementarity in general, then it is possible to extend Dwyer's hypothesis to a broader set of peptides: complementary peptides that bind to each other. In the latter case, one would expect to find homologous copies of the complementary peptide in the receptor. Thirteen peptides, 10 of which are not known to self-aggregate (amylin, ACTH, LHRH, angiotensin II, atrial natriuretic peptide, somatostatin, oxytocin, neurotensin, vasopressin, and substance P), and three that are known to self-aggregate (insulin, glucagon, and gastrin), were chosen. In addition to being self-aggregating, insulin and glucagon are also known to bind to each other, making them a mutually complementary pair. All possible combinations of the 13 peptides and the extracellular regions of their receptors were investigated using bioinformatic tools (a total of 325 combinations). Multiple, statistically significant homologies were found for insulin in the insulin receptor; insulin in the glucagon receptor; glucagon in the glucagon receptor; glucagon in the insulin receptor; and gastrin in gastrin binding protein and its receptor. Most of these homologies are in regions or sequences known to contribute to receptor binding of the respective hormone. These results suggest that the Dwyer hypothesis for receptor evolution may be generalizable beyond self-aggregating to complementary peptides. The evolution of receptors may have been driven by small molecule complementarity augmented by modular evolutionary processes that left a "molecular paleontology" that is still evident in the genome today. This "paleontology" may allow identification of peptide receptor sites.

Computational Biology↗

Bioinformatics in protein analysis.

The chapter gives an overview of bioinformatic techniques of importance in protein analysis. These include database searches, sequence comparisons and structural predictions. Links to useful World Wide Web (WWW) pages are given in relation to each topic. Databases with biological information are reviewed with emphasis on databases for nucleotide sequences (EMBL, GenBank, DDBJ), genomes, amino acid sequences (Swissprot, PIR, TrEMBL, GenePept), and three-dimensional structures (PDB). Integrated user interfaces for databases (SRS and Entrez) are described. An introduction to databases of sequence patterns and protein families is also given (Prosite, Pfam, Blocks). Furthermore, the chapter describes the widespread methods for sequence comparisons, FASTA and BLAST, and the corresponding WWW services. The techniques involving multiple sequence alignments are also reviewed: alignment creation with the Clustal programs, phylogenetic tree calculation with the Clustal or Phylip packages and tree display using Drawtree, njplot or phylo_win. Finally, the chapter also treats the issue of structural prediction. Different methods for secondary structure predictions are described (Chou-Fasman, Garnier-Osguthorpe-Robson, Predator, PHD). Techniques for predicting membrane proteins, antigenic sites and postranslational modifications are also reviewed.

Computational Biology↗

Application of knowledge information processing methods to biochemical engineering, biomedical and bioinformatics fields.

In biochemical and biomedical engineering fields there are a variety of phenomena with many complex chemical reactions, in which many genes and proteins affect transcription or enzyme activity of others. It is difficult to analyze and estimate many of these phenomena using conventional mathematical models. Recently some knowledge information processing methods, such as the artificial neural network (ANN), fuzzy reasoning, fuzzy neural network (FNN), fuzzy adaptive resonance theory (fuzzy ART) and the genetics algorithm (GA), were developed in the computer science field and have been applied to analysis in a variety of research fields. In this chapter, these methods will be briefly reviewed. Next, the application of these methods in the biochemical field will be introduced, instancing two examples in actual industrial processes. In addition, the application in the biomedical and bioinformatics field as another attractive field will be reviewed. Two examples are our research such as the prediction of prognosis for cancer patients from DNA microarray data using FNN and gene clustering for DNA microarray data using fuzzy ART.

Artificial Intelligence↗

Molecular modeling of protein structure and function: a bioinformatic approach.

This paper reports on the data/information structure of macromolecules as it extends beyond the three-dimensional conformation to include functional descriptors of biochemical (in vitro) and biological (in vivo) characteristics and as it contrasts with the limitations imposed by the data reduction and data classification techniques of traditional molecular modeling. Methodologies for structure-function representation are presented which are being incorporated within a knowledge-acquisition expert system. Examples of the bioinformatic approach are presented concerning macromolecular recognition by serine proteases and the use of Fourier transform-infrared (FT-IR) spectroscopy for structural assignment and analysis by a novel structure-perturbation approach.

Computer Simulation↗

The role of the pathologist as tissue refiner and data miner: the impact of functional genomics on the modern pathology laboratory and the critical roles of pathology informatics and bioinformatics.

This article provides an overview of how functional genomics is likely to impact on the pathology laboratory and highlights how informatics and tissue banking will greatly facilitate the molecular age of medicine. Important aspects of functional genomics in the post-genome era, including the roles of laser capture microdissection, DNA- and complementary DNA-based microarrays, proteomic methods, collaborative human tissue banking, tissue microarrays, and pathobioinformatics in the modern pathology laboratory are discussed. The role of mass spectroscopy in the analysis of RNA, DNA, and protein and its impact on the clinical laboratory, particularly in cost-effectiveness and time savings, are evaluated. This article explores how laboratory information systems (LISs) and the devices that feed them information may need to be modified to adapt to greater volumes of data for the new testing modalities that require understanding sophisticated fluorescence detection methods and image processing. Emerging genomic testing methods and their impact on pathology laboratory testing, especially in the area of molecular classification of neoplasms, are examined. The role of the tissue bank in the modern pathology laboratory as an archive of control normal tissues, as well as subsamples of the spectrum of progressive neoplastic states, is discussed in light of its critical importance to the molecular classification of cancer. Establishing a database that combines structured reports in pathology LISs and construction of tissue banking information systems will provide a rich resource for pathology departments. The article discusses a hypothetical resource, such as the Shared Tumor Expression Profiler, that would provide access to well-characterized tissue-based research resources for clinicians and researchers. Last, the article emphasizes how LISs can prepare for these changes, and how training pathologists in pathology informatics and bioinformatics (pathobioinformatics) is critical to ensure pathology's overall leadership role in the post-genome era.

Clinical Laboratory Techniques↗

Thrombin-like effect of an important green pit viper toxin, albolabrin: a bioinformatic study.

The green pit viper venom has a major effect on the hematological system. Clinical features of venomous snakebites vary from asymptomatic to fatal bleeding. The venom is found to have a thrombin-like effect in vitro. Here, the author performs a bioinformatic analysis on the green pit viper venom focusing on its thrombin-like effect. Sequence comparison between green pit viper venom, albolabrin and thrombin was performed. In addition, the author performed a search for other human proteins closely relating to the thrombin and created a multiple sequence alignment phylogenetic tree to present the family tree of the thrombin, albolabrin and those proteins recorded in the genomic database. In conclusion, the comparative sequence analysis between green pit viper venom and thrombin gives several identities. The reported relationship on the phylogenetic tree can match with the reported in vivo function of green pit viper. Explanations on the effect of green pit viper toxin on the hemostasis can be derived from this study. Furthermore, future researches based on the reported identities can be expected.

Amino Acid Sequence↗

Bioinformatic analysis of functional differences between the immunoproteasome and the constitutive proteasome.

Intracellular proteins are degraded largely by proteasomes. In cells stimulated with gamma interferon, the active proteasome subunits are replaced by "immuno" subunits that form immunoproteasomes. Phylogenetic analysis of the immunosubunits has revealed that they evolve faster than their constitutive counterparts. This suggests that the immunoproteasome has evolved a function that differs from that of the constitutive proteasome. Accumulating experimental degradation data demonstrate, indeed, that the specificity of the immunoproteasome and the constitutive proteasome differs. However, it has not yet been quantified how different the specificity of two forms of the proteasome are. The main question, which still lacks direct evidence, is whether the immunoproteasome generates more MHC ligands. Here we use bioinformatics tools to quantify these differences and show that the immunoproteasome is a more specific enzyme than the constitutive proteasome. Additionally, we predict the degradation of pathogen proteomes and find that the immunoproteasome generates peptides that are better ligands for MHC binding than peptides generated by the constitutive proteasome. Thus, our analysis provides evidence that the immunoproteasome has co-evolved with the major histocompatibility complex to optimize antigen presentation in vertebrate cells.

Animals↗

Bioinformatic discovery and initial characterisation of nine novel antimicrobial peptide genes in the chicken.

Antimicrobial peptides (AMPs) are essential components of innate immunity in a range of species from Drosophila to humans and are generally thought to act by disrupting the membrane integrity of microbes. In order to discover novel AMPs in the chicken, we have implemented a bioinformatic approach that involves the clustering of more than 420,000 chicken expressed sequence tags (ESTs). Similarity searching of proteins-predicted to be encoded by these EST clusters-for homology to known AMPs has resulted in the in silico identification of full-length sequences for seven novel gallinacins (Gal-4 to Gal-10), a novel cathelicidin and a novel liver-expressed antimicrobial peptide 2 (LEAP-2) in the chicken. Differential gene expression of these novel genes has been demonstrated across a panel of chicken tissues. An evolutionary analysis of the gallinacin family has detected sites-primarily in the mature AMP-that are under positive selection in these molecules. The functional implications of these results are discussed.

Amino Acid Sequence↗

Bioinformatic tools for DNA/protein sequence analysis, functional assignment of genes and protein classification.

The development of efficient DNA sequencing methods has led to the achievement of the DNA sequence of entire genomes from (to date) 55 prokaryotes, 5 eukaryotic organisms and 10 eukaryotic chromosomes. Thus, an enormous amount of DNA sequence data is available and even more will be forthcoming in the near future. Analysis of this overwhelming amount of data requires bioinformatic tools in order to identify genes that encode functional proteins or RNA. This is an important task, considering that even in the well-studied Escherichia coli more than 30% of the identified open reading frames are hypothetical genes. Future challenges of genome sequence analysis will include the understanding of gene regulation and metabolic pathway reconstruction including DNA chip technology, which holds tremendous potential for biomedicine and the biotechnological production of valuable compounds. The overwhelming volume of information often confuses scientists. This review intends to provide a guide to choosing the most efficient way to analyze a new sequence or to collect information on a gene or protein of interest by applying current publicly available databases and Web services. Recently developed tools that allow functional assignment of genes, mainly based on sequence similarity of the deduced amino acid sequence, using the currently available and increasing biological databases will be discussed.

Computational Biology↗

Serological identification and bioinformatics analysis of immunogenic antigens in multiple myeloma.

Identifying appropriate tumor antigens is critical to the development of successful specific cancer immunotherapy. Serological analysis of tumor antigens by a recombinant cDNA expression library (SEREX) allows the systematic cloning of tumor antigens recognized by the spontaneous autoantibody repertoire of cancer patients. We applied SEREX to the cDNA expression library of cell line HMy2, which led to the isolation of six known characterized genes and 12 novel genes. Known genes, including ring finger protein 167, KLF10, TPT1, p02 protein, cDNA FLJ46859 fis, and DNMT1, were related to the development of different tumors. Bioinformatics was performed to predict 12 novel MMSA (multiple myeloma special antigen) genes. The prediction of tumor antigens provides potential targets for the immunotherapy of patients with multiple myeloma (MM) and help in the understanding of carcinogenesis. Crude lysate ELISA methodology indicated that the optical density value of MMSA-3 and MMSA-7 were significantly higher in MM patients than in healthy donors. Furthermore, SYBR Green real-time PCR showed that MMSA-1 presented with a high number of copy messages in MM. In summary, the antigens identified in this study may be potential candidates for diagnosis and targets for immunotherapy in MM.

Antigens, Neoplasm↗

Bioinformatic and expression analysis of novel porcine beta-defensins.

Beta-defensins are a major group of mammalian antimicrobial peptides. Although more than 30 beta-defensins have been identified in humans, only one porcine beta-defensin has been reported. In this article we report the identification and initial characterization of 11 novel porcine beta-defensins (pBD). Using bioinformatic approaches, we screened 287,821 porcine expressed sequence tags for similarity of their predicted peptides to known human beta-defensins and identified full-length or partial sequences for the 11 novel pBDs. Similar to the previously identified pBD1, all of these peptides have a consensus beta-defensin motif. A differential expression pattern for these newly identified genes was found. For example, unlike most beta-defensins, pBD2 and pBD3 were expressed in bone marrow and in other lymphoid tissues including thymus, spleen, lymph nodes, duodenum, and liver. Including pBD2 and pBD3, six porcine beta-defensins were expressed in lung and skin. Several newly identified porcine beta-defensins, including pBD123, pBD125, and pBD129, were expressed in male reproductive tissues, including lobuli testis and some segments of the epididymis. Phylogenetic analysis indicates that in most cases the evolutionary relationship between individual porcine beta-defensins and their human orthologs is closer than the relationship among beta-defensins in the same species. These findings establish the existence of multiple porcine beta-defensins and suggest that the pig may be an ideal model for the characterization of beta-defensin diversity and function.

Amino Acid Sequence↗

Sequence analysis and bioinformatics analysis of chromosome 17q25 in familial moyamoya disease.

OBJECTS: The pathogenesis of moyamoya disease is still unknown. The present study aimed to find out the responsible genes that are located in the 17q25 locus. METHODS: Considering the function, we selected nine genes as candidates from a total of 65 genes identified in the 9-cM region of D17S785-D17S836 in chromosome 17q25, and performed sequence analysis on the DNA samples obtained from a pedigree of familial moyamoya disease, which showed a complete linkage to the region by a haplotype analysis. Also, we attempted to identify candidate genes that have not been known but might be functionally relevant to the disease among a total of 2,100 expressed sequence tag (EST) sequences using bioinformatics techniques. RESULTS AND CONCLUSION: The sequence analysis could detect no mutation in the nine genes. Nor could we identify a novel candidate gene by the EST analysis. Further studies using alternative approaches are warranted to clarify the pathogenesis of moyamoya disease.

Chromosomes, Human, Pair 17↗

Bioinformatic and molecular analysis of hydroxymethylbutenyl diphosphate synthase (GCPE) gene expression during carotenoid accumulation in ripening tomato fruit.

Carotenoids are plastidic isoprenoid pigments of great biological and biotechnological interest. The precursors for carotenoid production are synthesized through the recently elucidated methylerythritol phosphate (MEP) pathway. Here we have identified a tomato ( Lycopersicon esculentum Mill.) cDNA sequence encoding a full-length protein with homology to the MEP pathway enzyme hydroxymethylbutenyl 4-diphosphate synthase (HDS, also called GCPE). Comparison with other plant and bacterial HDS sequences showed that the plant enzymes contain a plastid-targeting N-terminal sequence and two highly conserved plant-specific domains in the mature protein with no homology to any other sequence in the databases. The ubiquitous distribution of HDS-encoding expressed sequence tags (ESTs) in the tomato collections suggests that the corresponding gene is likely expressed throughout the plant. The role of HDS in controlling the supply of precursors for carotenoid biosynthesis was estimated from the bioinformatic and molecular analysis of transcript abundance in different stages of fruit development. No significant changes in HDS gene expression were deduced from the statistical analysis of EST distribution during fruit ripening, when an active MEP pathway is required to support a massive accumulation of carotenoids. RNA blot experiments confirmed that similar transcript levels were present in both the wild-type and carotenoid-depleted yellow ripe ( r) mutant fruit independent of the stage of development and the carotenoid composition of the fruit. Together, our results are consistent with a non-limiting role for HDS in carotenoid biosynthesis during tomato fruit ripening.

Amino Acid Sequence↗