PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Bioinformatics -- a patenting view.

The use of bioinformatics in the biological sciences has brought about a change in the way that biological inventions can be protected by patent laws. Using approaches developed in the fields of computer science and business, patent applicants now seek to protect certain aspects of their inventions, which include software, methods of doing business and uses of information as well as more traditional biotechnological products and processes. These approaches are useful in resolving some of the difficulties now faced in prosecuting patent applications directed to biological inventions that are claimed in more conventional terms.

Algorithms↗

Can we integrate bioinformatics data on the Internet?

The NETTAB (Network Tools and Applications in Biology) 2001 Workshop entitled 'CORBA and XML: towards a bioinformatics-integrated network environment' was held at the Advanced Biotechnology Centre, Genoa, Italy, 17-18 May 2001.

Computational Biology↗

Bioinformatics methods for the analysis of expression arrays: data clustering and information extraction.

Expression arrays facilitate the monitoring of changes in the expression patterns of large collections of genes. The analysis of expression array data has become a computationally-intensive task that requires the development of bioinformatics technology for a number of key stages in the process, such as image analysis, database storage, gene clustering and information extraction. Here, we review the current trends in each of these areas, with particular emphasis on the development of the related technology being carried out within our groups.

Abstracting and Indexing↗

The TGF-beta--Smad network: introducing bioinformatic tools.

The TGF-beta superfamily is an important class of intercellular signalling molecule, including TGF-beta and bone morphogenetic proteins. Intracellular signalling cascades triggered by these molecules eventually activate transcription factors of the Smad family, which then regulate expression of their respective target genes. This article will discuss the TGF-beta--Smad signalling networks and how these processes are represented in databases of signal transduction and transcription control mechanisms. These databases can provide a well-structured overview of the subject and a basis for advanced bioinformatics analyses to interpret the function of genomic sequences or to analyse signalling networks.

Animals↗

Mapping cross-clade HIV-1 vaccine epitopes using a bioinformatics approach.

UNLABELLED: The genomic variability of HIV viruses circulating in different regions of the world has impeded the development of a globally relevant HIV vaccine. Broadly conserved HIV-1 cytotoxic T cell (CTL) epitopes were identified by screening protein sequences in the Los Alamos National Laboratory (LANL) HIV sequence database with a sequence parsing and matching algorithm (Conservatrix). Putative HIV-1 CTL epitopes were selected from this list using the epitope prediction tool EpiMatrix. METHODS: One hundred peptides representing putative HLA A*0201, HLA A*1101, HLA A*0301, and HLA B*07 ligands conserved in many isolates of HIV-1 were synthesized. Seventy-five HLA A*0201, HLA A*1101 and HLA B*07 peptides were incubated with transport associated protein (TAP)-deficient T2 cells transfected with the gene for the corresponding human HLA molecule (HLA A*0201, HLA A*1101, and HLA B*07). Binding and stabilization of peptide-HLA complexes on the surface of the T2 cells was measured by FACS. T cell responses to the entire set of 100 peptides (HLA A*0201, HLA A*1101, HLA A*0301, and HLA B*07) were measured in ELIspot assays using PBMC from healthy HIV-1 infected subjects who possessed a matching HLA allele. RESULTS: Fifty-seven (76%) of the 75 peptides tested in binding studies, including all (three of three) of the control (published) ligands bound to the T2 cells expressing the corresponding MHC molecule. Forty-three of the 100 peptides (43%) including all (four of four) of the control (published) epitopes tested in ELIspot assays stimulated gamma-interferon release. Thirty-one of these 43 epitopes are novel, highly conserved HIV-1 epitopes. EpiMatrix predicted and assays confirmed MHC-restriction by more than one HLA allele for nine of the 43 novel epitopes; of these epitopes five were recognized in the context of MHC "supertypes" and four were promiscuous epitopes. CONCLUSION: Epitopes identified using this approach were conserved in a broad range of HIV-1 sequences derived from isolates obtained in Latin America, Africa, Asia, the Pacific Islands, Europe and the US. The successful identification of cross-clade epitopes by this bioinformatics approach may accelerate the development of a globally relevant HIV-1 vaccine.

AIDS Vaccines↗

A bioinformatics based approach to discover small RNA genes in the Escherichia coli genome.

The recent explosion in available bacterial genome sequences has initiated the need to improve an ability to annotate important sequence and structural elements in a fast, efficient and accurate manner. In particular, small non-coding RNAs (sRNAs) have been difficult to predict. The sRNAs play an important number of structural, catalytic and regulatory roles in the cell. Although a few groups have recently published prediction methods for annotating sRNAs in bacterial genome, much remains to be done in this field. Toward the goal of developing an efficient method for predicting unknown sRNA genes in the completed Escherichia coli genome, we adopted a bioinformatics approach to search for DNA regions that contain a sigma70 promoter within a short distance of a rho-independent terminator. Among a total of 227 candidate sRNA genes initially identified, 32 were previously described sRNAs, orphan tRNAs, and partial tRNA and rRNA operons. Fifty-one are mRNAs genes encoding annotated extremely small open reading frames (ORFs) following an acceptable ribosome binding site. One hundred forty-four are potentially novel non-translatable sRNA genes. Using total RNA isolated from E. coli MG1655 cells grown under four different conditions, we verified transcripts of some of the genes by Northern hybridization. Here we summarize our data and discuss the rules and advantages/disadvantages of using this approach in annotating sRNA genes on bacterial genomes.

Base Sequence↗

MECP2 gene mutation analysis in the British and Italian Rett Syndrome patients: hot spot map of the most recurrent mutations and bioinformatic analysis of a new MECP2 conserved region.

Rett syndrome (RTT) is an X-linked dominant neurological disorder, which appears to be the most common genetic cause of profound combined intellectual and physical disability in Caucasian females. This syndrome has been associated with mutations of the MECP2 gene, a transcriptional repressor of unknown target genes. We report a detailed mutational analysis of a large cohort of RTT patients from the UK and Italy. This study has permitted us to produce a hot spot map of the mutations identified. Bioinformatic analysis of the mutations, taking advantage of structural and evolutionary data, leads us to postulate the existence of a new functional domain in the MeCP2 protein, conserved among brain-specific regulatory factors.

Adolescent↗

Combining structural and bioinformatics methods for the analysis of functionally important residues in DNA glycosylases.

An essential function of DNA glycosylases is the recognition and excision of damaged bases in DNA, thereby preserving genomic integrity. Lesion recognition is a multistep process, which is only partially revealed by structural analysis of the catalytically competent complex. The functional role of additional residues can be predicted by combining structural data with analysis of amino acid conservation. The following postulate underlies this approach: if a family or superfamily can be broken into subgroups with different substrate specificities, residues highly conserved between these subgroups represent those important for enzyme catalysis and structure maintenance while residues highly conserved within a subgroup but not between subgroups represent residues important for substrate specificity. We review the bioinformatics approach used for this quantitative analysis and describe its application to the Nth superfamily and Fpg family of DNA glycosylases. These results serve as a starting point in planning site-directed mutagenesis experiments to elucidate the functional role of similar and dissimilar residues in DNA repair and other proteins.

Amino Acid Sequence↗

PTEN and myotubularin phosphoinositide phosphatases: bringing bioinformatics to the lab bench.

Phosphoinositides play an integral role in a diverse array of cellular signaling processes. Although considerable effort has been directed toward characterizing the kinases that produce inositol lipid second messengers, the study of phosphatases that oppose these kinases remains limited. Current research is focused on the identification of novel lipid phosphatases such as PTEN and myotubularin, their physiologic substrates, signaling pathways and links to human diseases. The use of bioinformatics in conjunction with genetic analyses in model organisms will be essential in elucidating the roles of these enzymes in regulating phosphoinositide-mediated cellular signaling.

Amino Acid Sequence↗

Bioinformatics: from genome data to biological knowledge.

Recently, molecular biologists have sequenced about a dozen bacterial genomes and the first eukaryotic genome. We can now obtain answers to detailed questions about the complete set of genes of an organism. Bioinformatics methods are increasingly used for attaching biological knowledge to long lists of genes, assigning genes to biological pathways, comparing the gene sets of different species, identifying specificity factors, and describing sets of highly conserved proteins common to all domains of life. Substantial progress has recently been made in the availability of primary and added-value databases, in the development of algorithms and of network information services for genome analysis. The pharmaceutical industry has greatly benefited from the accumulation of sequence data through the identification of targets and candidates for the development of drugs, vaccines, diagnostic markers and therapeutic proteins.

Databases, Factual↗

The current excitement in bioinformatics-analysis of whole-genome expression data: how does it relate to protein structure and function?

Whole-genome expression profiles provide a rich new data-trove for bioinformatics. Initial analyses of the profiles have included clustering and cross-referencing to 'external' information on protein structure and function. Expression profile clusters do relate to protein function, but the correlation is not perfect, with the discrepancies partially resulting from the difficulty in consistently defining function. Other attributes of proteins can also be related to expression-in particular, structure and localization-and sometimes show a clearer relationship than function.

Computational Biology↗

Deriving folds of macromolecular complexes through electron cryomicroscopy and bioinformatics approaches.

Intermediate-resolution (7-9A) structures of large macromolecular complexes can be obtained by electron cryomicroscopy. This structural information, combined with bioinformatics data for the individual protein components or domains, can lead to a fold model for the entire complex. Such approaches have been demonstrated with the 6.8 A structure of the rice dwarf virus to derive models for the major capsid shell proteins.

Amino Acid Sequence↗

Bioinformatics tools for identifying class I-restricted epitopes.

The lack of simple methods to identify relevant T-cell epitopes, the high mutation rate of many pathogens, and restriction of T-cell response to epitopes due to human lymphocyte antigen (HLA) polymorphism have significantly hindered the development of cytotoxic T-lymphocyte (CTL) epitope-based or "epitope-driven" vaccines. Previously, CTL epitopes were mapped using large arrays of overlapping synthetic peptides. The large number of protein sequences available for mapping is now making this method prohibitively expensive and time-consuming. Bioinformatics tools such as EpiMatrix and Conservatrix, which search for unique or multi-HLA-restricted (promiscuous) T-cell epitopes and identify epitopes that are conserved across variant strains of the same pathogen, accelerate epitope mapping. These tools offer a significant advantage over other methods of epitope selection because high-throughput screening can be performed in silico, followed by confirmatory studies in vitro. CTL epitopes discovered using these tools might be used to develop novel vaccines and therapeutics for the prevention and treatment of infectious diseases such as human immunodeficiency virus, hepatitis C, tuberculosis, and some cancers.

Algorithms↗

Bioinformatic analysis of ClpS, a protein module involved in prokaryotic and eukaryotic protein degradation.

ClpS is a small protein, usually encoded immediately upstream of ClpA in the genomes of proteobacteria. Recent results show that it is a molecular adaptor for substrate recognition by ClpA in Escherichia coli. We analyzed ClpS by bioinformatic methods and found that ClpS homologs are also found in organisms that lack ClpA, such as actinobacteria, cyanobacteria, and plant chloroplasts. Furthermore, ClpS is homologous to a domain in the eukaryotic E3 ubiquitin ligase, N-recognin. This domain has previously been described as responsible for the recognition of type 2 N-end rule substrates. Despite very low levels of sequence similarity to proteins of known structure, there appears to be substantial structural similarity between ClpS and the C-terminal domain of ribosomal protein L7/12 (1CTF).

Adenosine Triphosphatases↗

Peptide libraries: at the crossroads of proteomics and bioinformatics.

Peptide libraries offer a valuable means for providing functional information regarding protein-modifying enzymes and protein interaction domains. Library approaches have become increasingly useful as high-throughput strategies for the analysis of large numbers of new proteins identified as a result of genome-sequencing efforts. Recent developments in the field have produced faster methods with broadened applicability. Crucially, new computational and biochemical tools have emerged that facilitate identification of interaction partners and substrates for proteins on the basis of their peptide selectivity profiles. Such combinations of proteomics-scale experimental approaches with bioinformatics tools hold great promise for the elucidation of protein interaction networks and signal transduction pathways in living cells.

Computational Biology↗

Progress in bioinformatics and the importance of being earnest.

In silico biology has gathered momentum as, worldwide, scientists have united in a common quest to sequence, store and analyse complete genomes. This year, a pivotal achievement of this cooperative endeavour was realised in the release of a public draft of the human genome, and with it the promises to improve our understanding of diverse aspects of biology and to yield a healthier future with safe personalized medicines. Key to these goals will be the need to elucidate and characterise the genes and gene products encoded not just in the human genome, but in many genomes. These tasks are underpinned by the concepts and processes of genome and gene/protein evolution, regulation of gene expression, mechanisms of protein folding, the manifestation of protein function, and so on, all of which must be understood in the context of complex, dynamic biological systems. Our use of computers to model such concepts and systems must be placed in the context of the current limits of our understanding of them:- it is important to recognise, for example, that we don't have a common understanding either of what constitutes a gene or a protein function; we can't invariably say that a particular sequence or fold has arisen via divergent or convergent evolution; and we don't fully understand the rules of protein folding. Accepting what we can't do in silico is essential in appreciating what we can do. Without this understanding, it is easy to be misled, as notions of what particular computational approaches can achieve are sometimes rather optimistic. There are valuable lessons to be learned here from the field of Artificial Intelligence, principal among which is the realisation that capturing and representing complex knowledge is time consuming, expensive and hard. Thus, we argue here that if bioinformatics is to tackle biological complexity in earnest, it would be wise to absorb the experience distilled from decades of artificial intelligence research, and to approach the road ahead with caution, rigour and pragmatism.

Artificial Intelligence↗

Bioinformatics in proteomics.

Several genome sequencing projects have recently been completed and the majority of human coding regions have been sequenced. In the next step many of the further studies will concentrate on proteins. Proteomics methods are essential for studying protein expression, activity, regulation and modifications. Bioinformatics is an integral part of proteomics research. The recent developments and applications in proteomics are discussed including mass spectrometry data analysis and interpretation, analysis and storage of the gel images to databases, gel comparison, and advanced methods to study e.g. protein co-expression, protein-protein interactions, as well as metabolic and cellular pathways. The significance of informatics in proteomics will gradually increase because of the advent of high-throughput methods relying on powerful data analysis.

Computational Biology↗

Bioinformatics approaches for the classification of G-protein-coupled receptors.

G-protein-coupled receptors are found abundantly in the human genome, and are the targets of numerous prescribed drugs. However, many receptors remain orphaned (i.e. with unknown ligand specificity), and others remain poorly characterised, with little structural information available. Consequently, there is often a gulf between sequence data and structural and functional knowledge of a receptor. Bioinformatics approaches may offer one approach to bridging this gap. In particular, protein family databases, which distil information from multiple sequence alignments into characteristic signatures, could be used to identify the families to which orphan receptors belong, and might facilitate discovery of novel motifs associated with ligand binding and G-protein-coupling.

Animals↗