PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

STRUCTURELAB: a heterogeneous bioinformatics system for RNA structure analysis.

STRUCTURELAB is a computational system that has been developed to permit the use of a broad array of approaches for the analysis of the structure of RNA. The goal of the development is to provide a large set of tools that can be well integrated with experimental biology to aid in the process of the determination of the underlying structure of RNA sequences. The approach taken views the structure determination problem as one of dealing with a database of many computationally generated structures and provides the capability to analyze this data set from different perspectives. Many algorithms are integrated into one system that also utilizes a heterogeneous computing approach permitting the use of several computer architectures to help solve the posed problems. These different computational platforms make it relatively easy to incorporate currently existing programs as well as newly developed algorithms and to best match these algorithms to the appropriate hardware. The system has been written in Common Lisp running on SUN or SGI Unix workstations, and it utilizes a network of participating machines defined in reconfigurable tables. A window-based interface makes this heterogeneous environment as transparent to the user as possible.

Algorithms↗

Applications of bioinformatics and computational biology to influenza surveillance and vaccine strain selection.

In recent years, collaborations often between mathematical and computational biologists and scientists in the World Health Organization (WHO) global influenza surveillance network, have resulted in a number of mathematical and computational advances including: increasing the resolution at which antigenic surveillance data can be analyzed, providing methods for genetic analysis and prediction, and an increased understanding of the determinants of repeated influenza vaccination. These advances increase the information extracted from influenza surveillance and increase the quantitative data available for the vaccine strain selection process. This mathematical and computational work is possible because of the wealth of information collected over many years by the WHO global influenza surveillance network, and further advances will be greatly facilitated by implementation of the proposed strengthening of virological and epidemiological surveillance in the WHO global agenda on influenza surveillance and control.

Computational Biology↗

Annealing function of GroEL: structural and bioinformatic analysis.

The Escherichia coli chaperonin system, GroEL-GroES, facilitates folding of substrate proteins (SPs) that are otherwise destined to aggregate. The iterative annealing mechanism suggests that the allostery-driven GroEL transitions leading to changes in the microenvironment of the SP constitutes the annealing action of chaperonins. To describe the molecular basis for the changes in the nature of SP-GroEL interactions we use the crystal structures of GroEL (T state), GroEL-ATP (R state) and the GroEL-GroES-(ADP)(7) (R" state) complex to determine the residue-specific changes in the accessible surface area and the number of tertiary contacts as a result of the T-->R-->R" transitions. We find large changes in the accessible area in many residues in the apical domain, but relatively smaller changes are associated with residues in the equatorial domain. In the course of the T-->R transition the microenvironment of the SP changes which suggests that GroEL is an annealing machine even without GroES. This is reflected in the exposure of Glu386 which loses six contacts in the T-->R transition. We also evaluate the conservation of residues that participate in the various chaperonin functions. Multiple sequence alignments and chemical sequence entropy calculations reveal that, to a large extent, only the chemical identities and not the residues themselves important for the nominal functions (peptide binding, nucleotide binding, GroES and substrate protein release) are strongly conserved. Using chemical sequence entropy, which is computed by classifying aminoacids into four types (hydrophobic, polar, positively charged and negatively charged) we make several new predictions that are relevant for peptide binding and annealing function of GroEL. We identify a number of conserved peptide binding sites in the apical domain which coincide with those found in the 1.7 A crystal structure of 'mini-chaperone' complexed with the N-terminal tag. Correlated mutations in the HSP60 family, that might control allostery in GroEL, are also strongly conserved. Most importantly, we find that charged solvent-exposed residues in the T state (Lys 226, Glu 252 and Asp 253) are strongly conserved. This leads to the prediction that mutating these residues, that control the annealing function of the SP, can decrease the efficacy of the chaperonin function.

Algorithms↗

Overlapping translation of nucleic acid sequences for bioinformatics applications.

SUMMARY: An alternative method to TblastX has been developed. Nucleic acids in database and query sequences were translated into overlapping protein-like sequences (overlappingly translated sequences or OTSs) before searching with BlastP. Thus, each nucleic acid sequences is represented by a single 'protein like' sequence instead of three 'proteins' in different reading frames. The 3x3 comparison of TblastX is represented by a single comparison, giving faster results. Additional advantages are: (1) it can be more sensitive to detect weak sequence similarities than either blastN or TblastX; (2) codon redundancy is eliminated; (3) the sensitivity to single nucleotide polymorphism, mutation and sequencing errors is reduced; (4) it is insensitive to frame shifts. RESULTS: BlastP using OTS detected about two thirds of blastN and TblastX matches but discovered additional similarities. When blastN and TblastX against nucleic acids were compared to blastP against OTS, identical matches discovered by blastP were generally longer (602, respectively. 213 letters, p<0.01), had higher scores (748 respectively 460 bits, p<0.05) and lower E values (3.16E-20 vs. 1.17E+03, p<0.01) but the percentage identity was lower (25% respectively 61%, p<0.001). A qualitative evaluation with LALIGN showed an improvement of the visualization when OTS-s were used instead of nucleic acids. Many extensive sequence similarities became better visible, for example the repeating similarity between prion protein and human insulin gene micro-satellite, and the surprising similarity between the first part of prion protein coding region and the human pro-insulin (34.4% identity and additional 17.2% similarity through 238 residues, score >295 which is expected 4.6e-18 times by chance).

Amino Acid Sequence↗

Identification of two hERR2-related novel nuclear receptors utilizing bioinformatics and inverse PCR.

Identification of novel nuclear receptors based on the highly conserved DNA-binding domain (DBD) has previously depended mainly on low stringency hybridization of cDNA libraries and degenerate PCR. Establishment of the expressed sequence tag (EST) database in recent years has provided an alternative approach for the discovery of novel members of gene families. The rate-limiting step is the conversion of ESTs to full-length cDNA. This article describes the identification of two novel nuclear receptors (hERRbeta2 and hERRgamma2) related to human estrogen-receptor-related receptor 2 (hERR2) by mining the EST database and retrieving of full-length cDNA via inverse PCR on subdivided primary cDNA library pools. The deduced protein sequences of hERRbeta2 and hERRgamma2 contain 500 and 458 amino acid (aa) residues respectively. Sequence analysis revealed that hERRbeta2 and hERRgamma2 respectively share 95% and 77% overall aa sequence identity with hERR2. However, the extra C-terminal domain in hERRbeta2 and extra N-terminal domain in hERRgamma2 are not present in the closely related hERR2 or mouse ERR2 (mERR2). Extensive sequence verification revealed that hERR2 previously reported as a human gene is actually a rat gene, whereas hERRbeta2 is the true human ortholog of hERR2 and mERR2. Tissue distribution studies showed that hERRgamma2 was expressed in a broader panel of tissues at a higher level than hERRbeta2. hERRbeta2 was mapped to cytogenetic locus 14q24.3 approximately -14q31, a region containing multiple loci involved in genetic diseases, including Alzheimer and diabetes. hERRgamma2 was mapped to 1q32. Given the high sequence homology between hERRbeta2 and mERR2, the two receptors may have similar biological function in vivo.

Amino Acid Sequence↗

A bioinformatic approach to the identification of bacterial proteins interacting with Toll-interleukin 1 receptor-resistance (TIR) homology domains.

Members of the Toll-like receptor (TLR) family are currently under intense scrutiny for their role in the sampling and recognition of pathogens. It has already been reported that both vaccinia virus and Yersinia spp. express proteins that help them evade the TLR mediated immune response, acting through the Toll-interleukin-1 receptor-resistance (TIR) domain and leucine-rich repeat region of the host TLRs respectively. The TIR domain is involved in the dimerisation of the TLRs and their complexation with their adapter molecules. We tested here the hypothesis that bacteria have the ability to secrete proteins containing similar motifs to the intracellular TIR domains that are involved in the TIR-TIR interaction necessary for the subsequent signal transmission. Based upon their sequence homology, proteins expressing TIRs have been divided into three sub-classes, based around the TLRs, the TLR adapter proteins, and the interleukin-1 and -18 adapter proteins. The highly conserved regions from these separate sub-families were then used to identify similar bacterial proteins. The bacterial proteins identified were then included in an iterative MEME-BLAST process to broaden the search. Tollip, a known TLR antagonist and adapter protein, was included in this investigation although it does not fit into any of the three sub-classes outlined above. If suitable bacterial proteins had been identified, it would signify that certain bacteria had evolved a mechanism to aid them in avoiding detection by the innate immune system acting through the TIR domains. At this stage one has to conclude that there is no evidence currently available suggesting such a mechanism, when using the strategy applied here.

Adaptor Proteins, Signal Transducing↗

Putting engineering back into protein engineering: bioinformatic approaches to catalyst design.

Complex multivariate engineering problems are commonplace and not unique to protein engineering. Mathematical and data-mining tools developed in other fields of engineering have now been applied to analyze sequence-activity relationships of peptides and proteins and to assist in the design of proteins and peptides with specified properties. Decreasing costs of DNA sequencing in conjunction with methods to quickly synthesize statistically representative sets of proteins allow modern heuristic statistics to be applied to protein engineering. This provides an alternative approach to expensive assays or unreliable high-throughput surrogate screens.

Algorithms↗

DNA repair: bioinformatics helps reverse methylation damage.

Recent work has uncovered a novel DNA repair enzyme: the AlkB protein of Escherichia coli, which oxidises the methyl groups of 1-methyladenine and 3-methylcytosine to hydroxymethyl moieties; the oxidised groups are subsequently released as formaldehyde, regenerating the unmodified bases.

Adenine↗

The use of bioinformatics for identifying class II-restricted T-cell epitopes.

An important step in the design of subunit vaccines is the identification of promiscuous T helper cell epitopes in sets of disease-specific gene products. Most of the epitope prediction models are based on HLA-II peptide binding, which constitutes a major bottleneck in the natural selection of epitopes. Here we describe a computer model, TEPITOPE, that enables the systematic prediction of promiscuous peptide ligands for a broad range of HLA binding specificity. We show how to apply the TEPITOPE prediction model to identify T-cell epitopes, and provide examples of its successful application in the context of oncology, allergy, and infectious and autoimmune diseases.

Computational Biology↗