PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Eukaryotic transcription factors in plastids--Bioinformatic assessment and implications for the evolution of gene expression machineries in plants.

The expression of genes in higher plant chloroplasts includes a complex transcriptional regulation which can be explained only in part with the action of the actually known components of the transcriptional machinery. This suggests the existence of still unknown important regulatory factors which influence chloroplast transcription. In order to test if such factors could exist we performed in silico analyses of Arabidopsis genes encoding putative transcription factors looking for putative N-terminal chloroplast transit peptides in the amino acid sequences. Our results suggest that 48 (and maybe up to 100) transcription factors of eukaryotic origin are likely to be imported into plastids. None of them has been described yet. This set of transcription factors highly expands the actually known regulation capacity of the chloroplast transcription machinery and provides a possible explanation for the complex initiation patterns of chloroplast transcripts. As consequence of a massive import of eukaryotic transcription factors a comprehensive reconstruction of the ancient prokaryotic gene expression machinery must be assumed resulting in a novel compatible combination of eukaryotic and prokaryotic protein components. In turn, the opposite process has been induced in the nucleus by the integration of prokaryotic components of the plastid ancestor via its loss of genes during endosymbiosis. Thus, a mutual exchange of regulatory factors, i.e. transcription factors occurred which resulted in the unique signalling network of today's plants. An evolutionary model of how this could have emerged during endosymbiosis in a timely coordinated manner is proposed.

Arabidopsis↗

Whole genome shotgun sequencing guided by bioinformatics pipelines--an optimized approach for an established technique.

While the sequencing of bacterial genomes has become a routine procedure at major sequencing centers, there are still a number of genome projects at small- or medium-size facilities. For these facilities a maximum of control over sequencing, assembling and finishing is essential. At the same time, facilities have to be able to co-operate at minimum costs for the overall project. We have established a pipeline for the distributed sequencing of Alcanivorax borkumensis SK2, Azoarcus sp. BH72, Clavibacter michiganensis subsp. michiganensis NCPPB382, Sorangium cellulosum So ce56 and Xanthomonas campestris pv. vesicatoria 85-10. Our pipeline relies on standard tools (e.g. PHRED/PHRAP, CAP3 and Consed/Autofinish) wherever possible, supplementing them with new tools (BioMake and BACCardI) to achieve the aims described above.

Algorithms↗

Bioinformatics support for high-throughput proteomics.

In the "post-genome" era, mass spectrometry (MS) has become an important method for the analysis of proteome data. The rapid advancement of this technique in combination with other methods used in proteomics results in an increasing number of high-throughput projects. This leads to an increasing amount of data that needs to be archived and analyzed. To cope with the need for automated data conversion, storage, and analysis in the field of proteomics, the open source system ProDB was developed. The system handles data conversion from different mass spectrometer software, automates data analysis, and allows the annotation of MS spectra (e.g. assign gene names, store data on protein modifications). The system is based on an extensible relational database to store the mass spectra together with the experimental setup. It also provides a graphical user interface (GUI) for managing the experimental steps which led to the MS data. Furthermore, it allows the integration of genome and proteome data. Data from an ongoing experiment was used to compare manual and automated analysis. First tests showed that the automation resulted in a significant saving of time. Furthermore, the quality and interpretability of the results was improved in all cases.

Algorithms↗

Bioinformatics database infrastructure for biotechnology research.

Many databases are available that provide valuable data resources for the biotechnological researcher. According to their core data, they can be divided into different types. Some databases provide primary data, like all published nucleotide sequences, others deal with protein sequences. In addition to these two basic types of databases, a huge number of more specialized resources are available, like databases about protein structures, protein identification, special features of genes and/or proteins, or certain organisms. Furthermore, some resources offer integrated views on different types of data, allowing the user to do easy customized queries over large datasets and to compare different types of data.

Animals↗

Structural bioinformatics prediction of membrane-binding proteins.

Membrane-binding peripheral proteins play important roles in many biological processes, including cell signaling and membrane trafficking. Unlike integral membrane proteins, these proteins bind the membrane mostly in a reversible manner. Since peripheral proteins do not have canonical transmembrane segments, it is difficult to identify them from their amino acid sequences. As a first step toward genome-scale identification of membrane-binding peripheral proteins, we built a kernel-based machine learning protocol. Key features of known membrane-binding proteins, including electrostatic properties and amino acid composition, were calculated from their amino acid sequences and tertiary structures, which were then incorporated into the support vector machine to perform the classification. A data set of 40 membrane-binding proteins and 230 non-membrane-binding proteins was used to construct and validate the protocol. Cross-validation and holdout evaluation of the protocol showed that the accuracy of the prediction reached up to 93.7% and 91.6%, respectively. The protocol was applied to the prediction of membrane-binding properties of four C2 domains from novel protein kinases C. Although these C2 domains have 50% sequence identity, only one of them was predicted to bind the membrane, which was verified experimentally with surface plasmon resonance analysis. These results suggest that our protocol can be used for predicting membrane-binding properties of a wide variety of modular domains and may be further extended to genome-scale identification of membrane-binding peripheral proteins.

Amino Acid Sequence↗

A bioinformatic analysis of the RAB genes of Trypanosoma brucei.

RAB proteins are small GTPases with vital roles in eukaryotic intracellular transport; orthologous RABs appear to fulfil similar functions in diverse organisms. Trypanosoma brucei spp., the causative organisms of Old World trypanosomiasis of humans and domestic animals, have extremely effective endocytic and exocytic mechanisms that are likely to be involved in maintenance of infection, making study of these systems of importance. Taking advantage of the essential completion of the T. brucei genome, we have re-examined the T. brucei RABs (TbRABs) so far described and identified a total of 16. BLAST searches and phylogenetic analysis show that nine of the TbRABs can confidently be assigned as orthologues or homologues of known RAB proteins from higher eukaryotes, and four more with reasonable probability. The core endocytic pathway is probably similar in complexity to yeast, whilst the early exocytic pathway appears to be more complex than in yeast. Two of the TbRAB family (RAB23 and 28) with clear mammalian orthologues appear to be unusual, and may be involved in nuclear processes and are described in more detail in an accompanying paper. Three TbRABs appear, however, to have no close homologues and may fulfil specialised functions in this organism. The availability of a complete set of TbRABs--which includes orthologues of the RABs responsible for control of the core of the endomembrane system (i.e. RAB1, 2, 4-7 and 11)--provides a first overview of the trafficking complexity that is present within a kinetoplastid parasite. Based on these homologies we suggest a systematic nomenclature for the TbRABs to reflect their functional homologies. This information is of importance both from the perspective of understanding the evolution and diversity of eukaryotic trafficking, but also in providing a framework by which to understand protein processing, trafficking, endocytosis and other related processes in these parasites.

Amino Acid Sequence↗

Update in bioinformatics. Toward a digital database of plant cell signalling networks: advantages, limitations and predictive aspects of the digital model.

The process of signal integration, which contributes to the regulation of multiple cellular activities, can be described in a digital language by a set of connected digital operations. In this article we delineate the basic concepts of cell signalling in the context of a logical description of information processing. Newly described instances of signal integration in plants are given as examples. The different advantages, limitations and predictive aspects of the digital modeling of signal transduction networks, as well as the minimal architecture of a computer database for plant signalling networks are discussed.

Computational Biology↗

Development of bioinformatic tools to support EST-sequencing, in silico- and microarray-based transcriptome profiling in mycorrhizal symbioses.

The great majority of terrestrial plants enters a beneficial arbuscular mycorrhiza (AM) or ectomycorrhiza (ECM) symbiosis with soil fungi. In the SPP 1084 "MolMyk: Molecular Basics of Mycorrhizal Symbioses", high-throughput EST-sequencing was performed to obtain snapshots of the plant and fungal transcriptome in mycorrhizal roots and in extraradical hyphae. To focus activities, the interactions between Medicago truncatula and Glomus intraradices as well as Populus tremula and Amanita muscaria were selected as models for AM and ECM symbioses, respectively. Together, almost, 20.000 expressed sequence tags (ESTs) were generated from different random and suppressive subtractive hybridization (SSH) cDNA libraries, providing a comprehensive overview of the mycorrhizal transcriptome. To automatically cluster and annotate EST-sequences, the BioMake and SAMS software tools were developed. In connection with the eNorthern software SteN, plant genes with a predicted mycorrhiza-induced expression were identified. To support experimental transcriptome profiling, macro- and microarray tools have been constructed for the two model mycorrhizae, based either on PCR-amplified cDNAs or 70mer oligonucleotides. These arrays were used to profile the transcriptome of AM and ECM roots under different conditions, and the data obtained were uploaded to the ArrayLIMS and EMMA databases that are designed to store and evaluate expression profiles from DNA arrays. Together, the EST- and transcriptome databases can be mined to identify candidate genes for targeted functional studies.

Computational Biology↗

Bioinformatic analyses of bacterial HPr kinase/phosphorylase homologues.

HPr kinase/phosphorylases (HprKs) regulate catabolite repression and sugar transport in Gram-positive bacteria by phosphorylating the small phosphotransferase system (PTS) protein HPr on a serine residue. We identified homologues of HprK in currently sequenced genomes and multiply aligned their sequences in order to perform phylogenetic and motif analyses. Seventy-eight homologues from bacteria and one from an archaeon comprise nine phylogenetic clusters. Some homologues come from bacteria whose genomes contain multiple highly divergent paralogues that cluster loosely together. Many of these proteins are truncated or show little or no identifiable similarity outside of the Walker A nucleotide binding domain. HprK homologues were identified in Gram-negative bacteria that appear to lack PTS permeases, suggesting modes of action and substrates that differ from those characterized in Gram-positive bacteria.

Amino Acid Sequence↗

Integrative bioinformatics for functional genome annotation: trawling for G protein-coupled receptors.

G protein-coupled receptors (GPCR) are amongst the best studied and most functionally diverse types of cell-surface protein. The importance of GPCRs as mediates or cell function and organismal developmental underlies their involvement in key physiological roles and their prominence as targets for pharmacological therapeutics. In this review, we highlight the requirement for integrated protocols which underline the different perspectives offered by different sequence analysis methods. BLAST and FastA offer broad brush strokes. Motif-based search methods add the fine detail. Structural modelling offers another perspective which allows us to elucidate the physicochemical properties that underlie ligand binding. Together, these different views provide a more informative and a more detailed picture of GPCR structure and function. Many GPCRs remain orphan receptors with no identified ligand, yet as computer-driven functional genomics starts to elaborate their functions, a new understanding of their roles in cell and developmental biology will follow.

Algorithms↗

Unique SARS-CoV protein nsp1: bioinformatics, biochemistry and potential effects on virulence.

Viruses have evolved a myriad of strategies for promoting viral replication, survival and spread. Sequence analysis of the Severe Acute Respiratory Syndrome Coronavirus (SARS-CoV) genome predicts several proteins that are unique to SARS-CoV. The search to understand the high virulence of SARS-CoV compared with related coronaviruses, which cause lesser respiratory illnesses, has recently focused on the unique nsp1 protein of SARS-CoV and suggests evolution of a possible new virulence mechanism in coronaviruses. The SARS-CoV nsp1 protein increases cellular RNA degradation and thus might facilitate SARS-CoV replication or block immune responses.

Animals↗

Mapping antigenic diversity and strain specificity of mumps virus: a bioinformatics approach.

Mumps is an acute infectious disease caused by mumps virus, a member of the family Paramyxoviridae. With the implementation of vaccination programs, mumps infection is under control. However, due to resurgence of mumps epidemics, there is a renewed interest in understanding the antigenic diversity of mumps virus. Hemagglutinin-neuraminidase (HN) is the major surface antigen and is known to elicit neutralizing antibodies. Mutational analysis of HN of wild-type and vaccine strains revealed that the hypervariable positions are distributed over the entire length with no detectable pattern. In the absence of experimentally derived 3D structure data, the structure of HN protein of mumps virus was predicted using homology modeling. Mutations mapped on the predicted structures were found to cluster on one of the surfaces. A predicted conformational epitope encompasses experimentally characterized epitopes suggesting that it is a major site for neutralization. These analyses provide rationale for strain specificity, antigenic diversity and varying efficacy of mumps vaccines.

Antigenic Variation↗

Bioinformatic characterization of the SynCAM family of immunoglobulin-like domain-containing adhesion molecules.

SynCAM 1 (synaptic cell adhesion molecule 1, alternatively named Tslc1 and nectin-like protein 3) belongs to the immunoglobulin superfamily and is an adhesion molecule that operates in a variety of important contexts. Exemplary are its roles in adhesion at synapses in the central nervous system and as tumor suppressor. Here, I describe a family of genes homologous to SynCAM 1 comprising four genes found solely in vertebrates. All SynCAM genes encode proteins with three immunoglobulin-like domains of the V-set, C1-set, and I-set subclasses. Comparison of genomic with cDNA sequences provides their exon-intron structure. Alternative splicing generates isoforms of SynCAM proteins, and diverse SynCAM 1 and 2 isoforms are created in an extracellular region rich in predicted O-glycosylation sites. Protein interaction motifs in the cytosolic sequence are highly conserved among all four SynCAM proteins, indicating their critical functional role. These findings aim to facilitate the understanding of SynCAM genes and provide the framework to examine the physiological functions of this family of vertebrate-specific adhesion molecules.

Alternative Splicing↗

Bioinformatics: searching the Net.

During the past 30 years, there has been an explosion in the volume of published medical information. As this volume has increased, so has the need for efficient methods for searching the data. MEDLINE, the primary medical database, is currently limited to abstracts of the medical literature. MEDLINE searches use AND/OR/NOT logical searching for keywords that have been assigned to each article and for textwords included in article abstracts. Recently, the complete text of some scientific journals, including figures and tables, has become accessible electronically. Keyword and textword searches can provide an overwhelming number of results. Search engines that use phrase searching, or searches that limit the number of words between two finds, improve the precision of search engines. The development of the Internet as a vehicle for worldwide communication, and the emergence of the World Wide Web (WWW) as a common vehicle for communication have made instantaneous access to much of the entire body of medical information an exciting possibility. There is more than one way to search the WWW for information. At the present time, two broad strategies have emerged for cataloging the WWW: directories and search engines. These allow more efficient searching of the WWW. Directories catalog WWW information by creating categories and subcategories of information and then publishing pointers to information within the category listings. Directories are analogous to yellow pages of the phone book. Search engines make no attempt to categorize information. They automatically scour the WWW looking for words and then automatically create an index of those words. When a specific search engine is used, its index is searched for a particular word. Usually, search engines are nonspecific and produce voluminous results. Use of AND/OR/NOT and "near" and "adjacent" search refinements greatly improve the results of a search. Search engines that limit their scope to specific sites, and metasearch sites that use multiple search engines optimized for specific types of searches have recently emerged. The distinctions between search engines and directory searches have blurred. Eventually, conceptual searching in which the computer searches for related ideas, without having to be given all the related keywords, may become a reality. This will free the user from having to learn specific rules about searching, allowing energies to be focused on results of the search, not the search itself.

Computational Biology↗