PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Bioinformatic and comparative localization of Rab proteins reveals functional insights into the uncharacterized GTPases Ypt10p and Ypt11p.

A striking characteristic of a Rab protein is its steady-state localization to the cytosolic surface of a particular subcellular membrane. In this study, we have undertaken a combined bioinformatic and experimental approach to examine the evolutionary conservation of Rab protein localization. A comprehensive primary sequence classification shows that 10 out of the 11 Rab proteins identified in the yeast (Saccharomyces cerevisiae) genome can be grouped within a major subclass, each comprising multiple Rab orthologs from diverse species. We compared the locations of individual yeast Rab proteins with their localizations following ectopic expression in mammalian cells. Our results suggest that green fluorescent protein-tagged Rab proteins maintain localizations across large evolutionary distances and that the major known player in the Rab localization pathway, mammalian Rab-GDI, is able to function in yeast. These findings enable us to provide insight into novel gene functions and classify the uncharacterized Rab proteins Ypt10p (YBR264C) as being involved in endocytic function and Ypt11p (YNL304W) as being localized to the endoplasmic reticulum, where we demonstrate it is required for organelle inheritance.

Animals↗

LegumeDB1 bioinformatics resource: comparative genomic analysis and novel cross-genera marker identification in lupin and pasture legume species.

The identification of markers in legume pasture crops, which can be associated with traits such as protein and lipid production, disease resistance, and reduced pod shattering, is generally accepted as an important strategy for improving the agronomic performance of these crops. It has been demonstrated that many quantitative trait loci (QTLs) identified in one species can be found in other plant species. Detailed legume comparative genomic analyses can characterize the genome organization between model legume species (e.g., Medicago truncatula, Lotus japonicus) and economically important crops such as soybean (Glycine max), pea (Pisum sativum), chickpea (Cicer arietinum), and lupin (Lupinus angustifolius), thereby identifying candidate gene markers that can be used to track QTLs in lupin and pasture legume breeding. LegumeDB is a Web-based bioinformatics resource for legume researchers. LegumeDB analysis of Medicago truncatula expressed sequence tags (ESTs) has identified novel simple sequence repeat (SSR) markers (16 tested), some of which have been putatively linked to symbiosome membrane proteins in root nodules and cell-wall proteins important in plant-pathogen defence mechanisms. These novel markers by preliminary PCR assays have been detected in Medicago truncatula and detected in at least one other legume species, Lotus japonicus, Glycine max, Cicer arietinum, and (or) Lupinus angustifolius (15/16 tested). Ongoing research has validated some of these markers to map them in a range of legume species that can then be used to compile composite genetic and physical maps. In this paper, we outline the features and capabilities of LegumeDB as an interactive application that provides legume genetic and physical comparative maps, and the efficient feature identification and annotation of the vast tracks of model legume sequences for convenient data integration and visualization. LegumeDB has been used to identify potential novel cross-genera polymorphic legume markers that map to agronomic traits, supporting the accelerated identification of molecular genetic factors underpinning important agronomic attributes in lupin.

Chromosome Mapping↗

Integration of bioinformatics and computational biology to understand protein-DNA recognition mechanism.

Transcription factors play essential role in the gene regulation in higher organisms, binding to multiple target sequences and regulating multiple genes in a complex manner. In order to decipher the mechanism of gene regulation, it is important to understand the molecular mechanism of protein-DNA recognition. Here we describe a strategy to approach this problem, using various methods in bioinformatics and computational biology. We have used a knowledge-based approach, utilizing rapidly increasing structural data of protein-DNA complexes, to derive empirical potential functions for the specific interactions between bases and amino acids as well as for DNA conformation, from the statistical analyses on the structural data. Then these statistical potentials are used to quantify the specificity of protein-DNA recognition. The quantification of specificity has enabled us to establish the structure-function analysis of transcription factors, such as the effects of binding cooperativity on target recognition. The method is also applied to real genome sequences, predicting potential target sites. We are also using computer simulations of protein-DNA interactions and DNA conformation in order to complement the empirical method. The integration of these approaches together will provide deeper insight into the mechanism of protein-DNA recognition and improve the target prediction of transcription factors.

Binding Sites↗

Bioinformatics tools for whole genomes.

The advent of whole-genome data resources--not only sequence but also other genome-scale data collections such as gene expression, protein interaction, and genetic variation--is having two marked, complementary effects on the relatively new discipline of bioinformatics. First, the veritable flood of data is creating a need and demand for new tools for dealing adequately with the deluge, and, second, the unprecedented extent, diversity, and impending completeness of the data sets are creating opportunities for new approaches to discovery based on computational methods.

Algorithms↗

Bioinformatic identification of novel early stress response genes in rodent models of lung injury.

Acute lung injury is a complex illness with a high mortality rate (>30%) and often requires the use of mechanical ventilatory support for respiratory failure. Mechanical ventilation can lead to clinical deterioration due to augmented lung injury in certain patients, suggesting the potential existence of genetic susceptibility to mechanical stretch (6, 48), the nature of which remains unclear. To identify genes affected by ventilator-induced lung injury (VILI), we utilized a bioinformatic-intense candidate gene approach and examined gene expression profiles from rodent VILI models (mouse and rat) using the oligonucleotide microarray platform. To increase statistical power of gene expression analysis, 2,769 mouse/rat orthologous genes identified on RG_U34A and MG_U74Av2 arrays were simultaneously analyzed by significance analysis of microarrays (SAM). This combined ortholog/SAM approach identified 41 up- and 7 downregulated VILI-related candidate genes, results validated by comparable expression levels obtained by either real-time or relative RT-PCR for 15 randomly selected genes. K-mean clustering of 48 VILI-related genes clustered several well-known VILI-associated genes (IL-6, plasminogen activator inhibitor type 1, CCL-2, cyclooxygenase-2) with a number of stress-related genes (Myc, Cyr61, Socs3). The only unannotated member of this cluster (n = 14) was RIKEN_1300002F13 EST, an ortholog of the stress-related Gene33/Mig-6 gene. The further evaluation of this candidate strongly suggested its involvement in development of VILI. We speculate that the ortholog-SAM approach is a useful, time- and resource-efficient tool for identification of candidate genes in a variety of complex disease models such as VILI.

Algorithms↗

Vitamin E deficiency and metabolic deficits in neuronal ceroid lipofuscinosis described by bioinformatics.

The mnd mouse, a model of neuronal ceroid lipofusinosis (NCL), has a profound vitamin E deficiency in sera and brain, associated with cerebral deterioration characteristic of NCL. In this study, the vitamin E deficiency is corrected using dietary supplementation. However, the histopathological features associated with NCL remained. With use of a bioinformatics approach based on high-resolution solid and solution state 1H-NMR spectroscopy and principal component analysis (PCA), the deficits associated with NCL are defined in terms of a metabolic phenotype. Although vitamin E supplementation reversed some of the metabolic abnormalities, in particular the concentration of phenylalanine in extracts of cerebral tissue, PCA demonstrated that metabolic deficits associated with NCL were greater than any effects produced from vitamin E supplementation. These deficits included increased glutamate and N-acetyl-L-aspartate and decreased creatine and glutamine concentrations in aqueous extracts of the cortex, as well as profound accumulation of lipid in intact cerebral tissue. This is discussed in terms of faulty production of mitochondrial-associated membranes, thought to be central to the deficits in mnd mice.

Animals↗

Metabolome, transcriptome, and bioinformatic cis-element analyses point to HNF-4 as a central regulator of gene expression during enterocyte differentiation.

DNA-binding transcription factors bind to promoters that carry their binding sites. Transcription factors therefore function as nodes in gene regulatory networks. In the present work we used a bioinformatic approach to search for transcription factors that might function as nodes in gene regulatory networks during the differentiation of the small intestinal epithelial cell. In addition we have searched for connections between transcription factors and the villus metabolome. Transcriptome data were generated from mouse small intestinal villus, crypt, and fetal intestinal epithelial cells. Metabolome data were generated from crypt and villus cells. Our results show that genes that are upregulated during fetal to adult and crypt to villus differentiation have an overrepresentation of potential hepatocyte nuclear factor (HNF)-4 binding sites in their promoters. Moreover, metabolome analyses by magic angle spinning (1)H nuclear magnetic resonance spectroscopy showed that the villus epithelial cells contain higher concentrations of lipid carbon chains than the crypt cells. These findings suggest a model where the HNF-4 transcription factor influences the villus metabolome by regulating genes that are involved in lipid metabolism. Our approach also identifies transcription factors of importance for crypt functions such as DNA replication (E2F) and stem cell maintenance (c-Myc).

Algorithms↗

The World-Wide Web: an interface between research and teaching in bioinformatics.

The rapid expansion occurring in World-Wide Web activity is beginning to make the concepts of 'global hypermedia' and 'universal document readership realistic objectives of the new revolution in information technology. One consequence of this increase in usage is that educators and students are becoming more aware of the diversity of the knowledge base which can be accessed via the Internet. Although computerised databases and information services have long played a key role in bioinformatics these same resources can also be used to provide core materials for teaching and learning. The large datasets and archives that have been compiled for biomedical research can be enhanced with the addition of a variety of multimedia elements (images, digital videos, animation etc.). The use of this digitally stored information in structured and self-directed learning environments is likely to increase as activity across World-Wide Web increases.

Information Systems↗

Postgenomics: Proteomics and Bioinformatics in Cancer Research.

Now that the human genome is completed, the characterization of the proteins encoded by the sequence remains a challenging task. The study of the complete protein complement of the genome, the "proteome," referred to as proteomics, will be essential if new therapeutic drugs and new disease biomarkers for early diagnosis are to be developed. Research efforts are already underway to develop the technology necessary to compare the specific protein profiles of diseased versus nondiseased states. These technologies provide a wealth of information and rapidly generate large quantities of data. Processing the large amounts of data will lead to useful predictive mathematical descriptions of biological systems which will permit rapid identification of novel therapeutic targets and identification of metabolic disorders. Here, we present an overview of the current status and future research approaches in defining the cancer cell's proteome in combination with different bioinformatics and computational biology tools toward a better understanding of health and disease.

Journal Article↗

Delineation, functional validation, and bioinformatic evaluation of gene expression in thyroid follicular carcinomas with the PAX8-PPARG translocation.

A subset of follicular thyroid carcinomas contains a balanced translocation, t(2;3)(q13;p25), that results in fusion of the paired box gene 8 (PAX8) and peroxisome proliferator-activated receptor gamma (PPARG) genes with concomitant expression of a PAX8-PPARgamma fusion protein, PPFP. PPFP is thought to contribute to neoplasia through a mechanism in which it acts as a dominant-negative inhibitor of wild-type PPARgamma. To better understand this type of follicular carcinoma, we generated global gene expression profiles using DNA microarrays of a cohort of follicular carcinomas along with other common thyroid tumors and used the data to derive a gene expression profile characteristic of PPFP-positive tumors. Transient transfection assays using promoters of four genes whose expression was highly associated with the translocation showed that each can be activated by PPFP. PPFP had unique transcriptional activities when compared with PAX8 or PPARgamma, although it had the potential to function in ways qualitatively similar to PAX8 or PPARgamma depending on the promoter and cellular environment. Bioinformatics analyses revealed that genes with increased expression in PPFP-positive follicular carcinomas include known PPAR target genes; genes involved in fatty acid, amino acid, and carbohydrate metabolism; micro-RNA target genes; and genes on chromosome 3p. These results have implications for the neoplastic mechanism of these follicular carcinomas.

Adenocarcinoma, Follicular↗

Chemosensitivity profile of cancer cell lines and identification of genes determining chemosensitivity by an integrated bioinformatical approach using cDNA arrays.

We have established a panel of 45 human cancer cell lines (JFCR-45) to explore genes that determine the chemosensitivity of these cell lines to anticancer drugs. JFCR-45 comprises cancer cell lines derived from tumors of three different organs: breast, liver, and stomach. The inclusion of cell lines derived from gastric and hepatic cancers is a major point of novelty of this study. We determined the concentration of 53 anticancer drugs that could induce 50% growth inhibition (GI50) in each cell line. Cluster analysis using the GI50s indicated that JFCR-45 could allow classification of the drugs based on their modes of action, which coincides with previous findings in NCI-60 and JFCR-39. We next investigated gene expression in JFCR-45 and developed an integrated database of chemosensitivity and gene expression in this panel of cell lines. We applied a correlation analysis between gene expression profiles and chemosensitivity profiles, which revealed many candidate genes related to the sensitivity of cancer cells to anticancer drugs. To identify genes that directly determine chemosensitivity, we further tested the ability of these candidate genes to alter sensitivity to anticancer drugs after individually overexpressing each gene in human fibrosarcoma HT1080. We observed that transfection of HT1080 cells with the HSPA1A and JUN genes actually enhanced the sensitivity to mitomycin C, suggesting the direct participation of these genes in mitomycin C sensitivity. These results suggest that an integrated bioinformatical approach using chemosensitivity and gene expression profiling is useful for the identification of genes determining chemosensitivity of cancer cells.

Antineoplastic Agents↗

Using bioinformatics and genome analysis for new therapeutic interventions.

The genome era provides two sources of knowledge to investigators whose goal is to discover new cancer therapies: first, information on the 20,000 to 40,000 genes that comprise the human genome, the proteins they encode, and the variation in these genes and proteins in human populations that place individuals at risk or that occur in disease; second, genome-wide analysis of cancer cells and tissues leads to the identification of new drug targets and the design of new therapeutic interventions. Using genome resources requires the storage and analysis of large amounts of diverse information on genetic variation, gene and protein functions, and interactions in regulatory processes and biochemical pathways. Cancer bioinformatics deals with organizing and analyzing the data so that important trends and patterns can be identified. Specific gene and protein targets on which cancer cells depend can be identified. Therapeutic agents directed against these targets can then be developed and evaluated. Finally, molecular and genetic variation within a population may become the basis of individualized treatment.

Computational Biology↗

Bioinformatic methods for allergenicity assessment using a comprehensive allergen database.

BACKGROUND: A principal aim of the safety assessment of genetically modified crops is to prevent the introduction of known or clinically cross-reactive allergens. Current bioinformatic tools and a database of allergens and gliadins were tested for the ability to identify potential allergens by analyzing 6 Bacillus thuringiensis insecticidal proteins, 3 common non-allergenic food proteins and 50 randomly selected corn (Zea mays) proteins. METHODS: Protein sequences were compared to allergens using the FASTA algorithm and by searching for matches of 6, 7 or 8 contiguous identical amino acids. RESULTS: No significant sequence similarities or matches of 8 contiguous amino acids were found with the B. thuringiensis or food proteins. Surprisingly, 41 of 50 corn proteins matched at least one allergen with 6 contiguous identical amino acids. Only 7 of 50 corn proteins matched an allergen with 8 contiguous identical amino acids. When assessed for overall structural similarity to allergens, these 7 plus 2 additional corn proteins shared >or=35% identity in an overlap of >or=80 amino acids, but only 6 of the 7 were similar across the length of the protein, or shared >50% identity to an allergen. CONCLUSIONS: An evaluation of a protein by the FASTA algorithm is the most predictive of a clinically relevant cross-reactive allergen. An additional search for matches of 8 amino acids may provide an added margin of safety when assessing the potential allergenicity of a protein, but a search with a 6-amino-acid window produces many random, irrelevant matches.

Algorithms↗

Use of bioinformatics to predict a function for the GS element in Mycobacterium avium subspecies paratuberculosis.

Mycobacterium avium subsp. Paratuberculosis (MAP) is a member of the Mycobacterium avium complex (MAC) and causes the inflammatory bowel disease, Johne's disease, in livestock. MAP has also been implicated as the causative agent of a similar disease, Crohn's disease, in humans. One of three major genetic differences between MAP and non-pathogenic MAC is the 6496-bp GS element. Based on the output from freely available protein sequence and structural bioinformatics tools, and the close homology of GS genes with the SER2 region of the closely related Mycobacterium avium subsp. Avium (MAA), we predict that GS encoded enzymes are involved in the biosynthesis of GDP-fucose, and the addition to, and modification of fucose on, the oligosaccharide moiety of GPL. GPL is a major constituent of the cell wall of the MAC and has immunomodulatory properties. Therefore, the enzymes involved in its synthesis may provide novel drug targets against MAP and other pathogenic MAC members.

Amino Acid Sequence↗

The bioinformatics of psychosocial genomics in alternative and complementary medicine.

The bioinformatics of alternative and complementary medicine is outlined in 3 hypotheses that extend the molecular-genomic revolution initiated by Watson and Crick 50 years ago to include psychology in the new discipline of psychosocial and cultural genomics. Stress-induced changes in the alternative splicing of genes demonstrate how psychosomatic stress in humans modulates activity-dependent gene expression, protein formation, physiological function, and psychological experience. The molecular messengers generated by stress, injury, and disease can activate immediate early genes within stem cells so that they then signal the target genes required to synthesize the proteins that will transform (differentiate) stem cells into mature well-functioning tissues. Such activity-dependent gene expression and its consequent activity-dependent neurogenesis and stem cell healing is proposed as the molecular-genomic-cellular basis of rehabilitative medicine, physical, and occupational therapy as well as the many alternative and complementary approaches to mind-body healing. The therapeutic replaying of enriching life experiences that evoke the novelty-numinosum-neurogenesis effect during creative moments of art, music, dance, drama, humor, literature, poetry, and spirituality, as well as cultural rituals of life transitions (birth, puberty, marriage, illness, healing, and death) can optimize consciousness, personal relationships, and healing in a manner that has much in common with the psychogenomic foundations of naturalistic and complementary medicine. The entire history of alternative and complementary approaches to healing is consistent with this new neuroscience world view about the role of psychological arousal and fascination in modulating gene expression, neurogenesis, and healing via the psychosocial and cultural rites of human societies.

Adaptation, Physiological↗

Bioinformatic analyses of the bacterial L-ascorbate phosphotransferase system permease family.

The tripartite L-ascorbate permease of Escherichia coli is the first functionally characterized member of a large family of enzyme II complexes (SgaTBA, encoding enzymes IIC, IIB and IIA) of the bacterial phosphotransferase system (PTS). We here report bioinformatic analyses of these proteins. Forty-five homologous systems from a wide variety of bacteria were identified, but no homologues were found in archaea or eukaryotes. These systems fell into five structural types: (1) IIC, IIB and IIA are encoded by distinct genes; (2) IIC and IIB are encoded by distinct genes, but the IIA-encoding gene is absent; (3) IIC and IIB are encoded by a fused gene, but IIA is a distinct gene product; (4) IIA and IIB are fused, but IIC is encoded by a distinct gene, and (5) IIC and IIB are encoded by distinct genes, but IIA is fused to a transcriptional regulator. Phylogenetic analyses revealed that gene fusion/splicing events have occurred repeatedly during the evolutionary divergence of family members, although no evidence for shuffling of constituents between systems was obtained. The SgaTBA family proved to be distantly related to the GatCBA family of PTS permeases, and this family was also analyzed. In contrast to the SgaTBA family, no gene splicing/fusion has occurred during the evolutionary divergence of GatCBA family members as each domain is always encoded by a distinct gene. However, GatC homologues were identified in organisms that lack other PTS proteins, suggesting a transport mechanism not coupled to substrate phosphorylation. Topological analyses suggest that in contrast to all other PTS permeases, IIC proteins of the Sga and Gat families exhibit 12 transmembrane alpha-helical segments and are distantly related to secondary carriers. Like many secondary carriers, GatC (IIC) homologues could be shown to have arisen by an ancient intragenic duplication event. These results suggest that the Sga and Gat families of PTS permeases comprise a small superfamily in which the transmembrane IIC domains evolved independently of all other known PTS permeases.

Amino Acid Transport Systems↗

Mycobacterial gene cloning and expression, comparative genomics, bioinformatics and proteomics in relation to the development of new vaccines and diagnostic reagents.

Recent advances in molecular and genomic techniques have facilitated research on several aspects of mycobacteriology, such as diagnosis and the identification of new vaccines and therapeutic targets for various diseases, including tuberculosis. The aim of this review was to analyze the implications of advances in molecular and genomic techniques on the development of new vaccines for tuberculosis as well as immunological reagents to diagnose the disease. Gene cloning and expression, DNA and protein sequencing, polymerase chain reaction, comparative genomics, bioinformatics, proteomics and DNA and peptide synthesis coupled with the application of cellular immunology techniques have led to the identification of several antigens of Mycobacterium tuberculosis, which have potential for diagnosis and vaccine applications. For example, cross-reactive mycobacterial antigens like heat shock proteins, MTB32 and MTB39, have been identified as new vaccine candidates, and antigens encoded by M. tuberculosis-specific genomic regions as new reagents for diagnosis.

Antigens, Bacterial↗

Development of bioinformatics resources for display and analysis of copy number and other structural variants in the human genome.

The discovery of an abundance of copy number variants (CNVs; gains and losses of DNA sequences >1 kb) and other structural variants in the human genome is influencing the way research and diagnostic analyses are being designed and interpreted. As such, comprehensive databases with the most relevant information will be critical to fully understand the results and have impact in a diverse range of disciplines ranging from molecular biology to clinical genetics. Here, we describe the development of bioinformatics resources to facilitate these studies. The Database of Genomic Variants (http://projects.tcag.ca/variation/) is a comprehensive catalogue of structural variation in the human genome. The database currently contains 1,267 regions reported to contain copy number variation or inversions in apparently healthy human cases. We describe the current contents of the database and how it can serve as a resource for interpretation of array comparative genomic hybridization (array CGH) and other DNA copy imbalance data. We also present the structure of the database, which was built using a new data modeling methodology termed Cross-Referenced Tables (XRT). This is a generic and easy-to-use platform, which is strong in handling textual data and complex relationships. Web-based presentation tools have been built allowing publication of XRT data to the web immediately along with rapid sharing of files with other databases and genome browsers. We also describe a novel tool named eFISH (electronic fluorescence in situ hybridization) (http://projects.tcag.ca/efish/), a BLAST-based program that was developed to facilitate the choice of appropriate clones for FISH and CGH experiments, as well as interpretation of results in which genomic DNA probes are used in hybridization-based experiments.

Algorithms↗