PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Bioinformatic approaches for accurate assessment of A-to-I editing in complete transcriptomes.

A-to-I RNA editing is an RNA modification that alters the RNA sequence relative to the its genomic blueprint. It is catalyzed by double-stranded RNA-specific adenosine deaminase (ADAR) enzymes, and contributes to the complexity and diversification of the proteome. Advancement in the study of A-to-I RNA editing has been facilitated by computational approaches for accurate mapping and quantification of A-to-I RNA editing based on sequencing data. In this chapter we review some of the main computational approaches currently used, describe potential hurdles, challenges and pitfalls, and discuss possible ways to mitigate them.

RNA Editing↗

How bioinformatics can help reverse engineer human aging.

To study human aging is an enormous challenge. The complexity of the aging phenotype and the near impossibility of studying aging directly in humans oblige researchers to resort to models and extrapolations. Computational approaches offer a powerful set of tools to study human aging. In one direction we have data-mining methods, from comparative genomics to DNA microarrays, to retrieve information in large amounts of data. Afterwards, tools from systems biology to reverse engineering algorithms allow researchers to integrate different types of information to increase our knowledge about human aging. Computer methodologies will play a crucial role to reconstruct the genetic network of human aging and the associated regulatory mechanisms.

Aging↗

Evidence for lectin activity of a plant receptor-like protein kinase by application of neoglycoproteins and bioinformatic algorithms.

Detection of genes for putative receptor-like protein kinases, which contain an extracellular domain related to leguminous lectins, in plant genomes inspired the hypothesis that this part acts as sensor. Initial support for this concept came from proof for protein kinase activity. The next step, focusing on the protein of lombardy poplar (Populus nigra var. italica), is scrutiny for lectin activity. Consequently, we first pinpointed sets of high-scoring sequence pairs by extensive databank search. The calculations resulted in P-values in the range from 10(-14) to 10(-18) exclusively for leguminous lectins, the Pterocarpus angolensis agglutinin being front runner with P=3 x 10(-18) and thus most suitable template for modeling. The superimposition of the two folds gave notable similarity in the region responsible for binding carbohydrate and Ca(2+)/Mn(2+)-ions. Binding activity toward carbohydrates was detected by assaying a panel of (neo)glycoproteins as polyvalent probes, especially for alpha-l-rhamnose and glycans of asialofetuin. It was strictly dependent on Ca(2+)-ions, enhanced by Mn(2+)-ions and reached a K(D)-value of 34.3 nM for the neoglycoprotein with rhamnose as ligand. These results give further research direction to define physiological ligands, plant/bacterial rhamnose-containing saccharides and rhamnose-mimetic glycans or peptides being potential candidates.

Algorithms↗

Structural bioinformatics study of EPSP synthase from Mycobacterium tuberculosis.

The shikimate pathway is an attractive target for herbicides and antimicrobial agent development because it is essential in algae, higher plants, bacteria, and fungi, but absent from mammals. Homologues to enzymes in the shikimate pathway have been identified in the genome sequence of Mycobacterium tuberculosis. Among them, the EPSP synthase was proposed to be present by sequence homology. Accordingly, in order to pave the way for structural and functional efforts towards anti-mycobacterial agent development, here we describe the molecular modeling of 5-enolpyruvylshikimate-3-phosphate (EPSP) synthase isolated from M. tuberculosis that should provide a structural framework on which the design of specific inhibitors may be based on. Significant differences in the relative orientation of the domains in the two models result in "open" and "closed" conformations. The possible relevance of this structural transition in the ligand biding is discussed.

3-Phosphoshikimate 1-Carboxyvinyltransferase↗

Structural bioinformatics study of PNP from Schistosoma mansoni.

The parasite Schistosoma mansoni lacks the de novo pathway for purine biosynthesis and depends on salvage pathways for its purine requirements. Schistosomiasis is endemic in 76 countries and territories and amongst the parasitic diseases ranks second after malaria in terms of social and economic impact and public health importance. The PNP is an attractive target for drug design and it has been submitted to extensive structure-based design. The atomic coordinates of the complex of human PNP with inosine were used as template for starting the modeling of PNP from S. mansoni complexed with inosine. Here we describe the model for the complex SmPNP-inosine and correlate the structure with differences in the affinity for inosine presented by human and S. mansoni PNPs.

Amino Acid Sequence↗

Expression analysis of members of the neuronal calcium sensor protein family: combining bioinformatics and Western blot analysis.

We have used in silico mining of public databases (NCBI UniGene and NCI SAGE Anatomic Viewer) as a tool to obtain the tissue distribution pattern of three members of the neuronal calcium sensor protein family, namely VILIP-1, hippocalcin, and NCS-1 in humans. The theoretical human mRNA expression profile of the calcium sensor proteins derived from expressed sequence tag (EST) and serial analysis of gene expression (SAGE) data was compared with expression data from human tissues obtained by Western blot analysis. Since the EST databank searches do not yet give comparable results for rat which is often used as model animal, we have also analyzed the protein expression in rat tissues. Similar to the human expression profile in rat tissues calcium sensor proteins are mainly detected in the nervous system, but the data consistently implicated the additional expression in peripheral tissues with remarkable differences between the calcium sensor proteins.

Animals↗

Characterization of histone (H1B) oxalate binding protein in experimental urolithiasis and bioinformatics approach to study its oxalate interaction.

The rat kidney H1 oxalate binding protein was isolated and purified. Oxalate binds exclusively with H1B fraction of H1 histone. Oxalate binding activity is inhibited by lysine group modifiers such as 4',4'-diisothiostilbene-2,2-disulfonic acid (DIDS) and pyridoxal phosphate and reduced in presence of ATP and ADP. RNA has no effect on oxalate binding activity of H1B whereas DNA inhibits oxalate binding activity. Equilibrium dialysis method showed that H1B oxalate binding protein has two binding sites for oxalate, one with high affinity, other with low affinity. Histone H1B was modeled in silico using Modeller8v1 software tool since experimental structure is not available. In silico interaction studies predict that histone H1B-oxalate interaction take place through lysine121, lysine139, and leucine68. H1B oxalate binding protein is found to be a promoter of calcium oxalate crystal (CaOx) growth. A 10% increase in the promoting activity is observed in hyperoxaluric rat kidney H1B. Interaction of H1B oxalate binding protein with CaOx crystals favors the formation of intertwined calcium oxalate dehydrate (COD) crystals as studied by light microscopy. Intertwined COD crystals and aggregates of COD crystals were more pronounced in the presence of hyperoxalauric H1B.

Amino Acid Sequence↗

Isolation, characterization, and bioinformatic analysis of calmodulin-binding protein cmbB reveals a novel tandem IP22 repeat common to many Dictyostelium and Mimivirus proteins.

A novel calmodulin-binding protein cmbB from Dictyostelium discoideum is encoded in a single gene. Northern analysis reveals two cmbB transcripts first detectable at 4 h during multicellular development. Western blotting detects an approximately 46.6 kDa protein. Sequence analysis and calmodulin-agarose binding studies identified a "classic" calcium-dependent calmodulin-binding domain (179IPKSLRSLFLGKGYNQPLEF198) but structural analyses suggest binding may not involve classic alpha-helical calmodulin-binding. The cmbB protein is comprised of tandem repeats of a newly identified IP22 motif ([I,L]Pxxhxxhxhxxxhxxxhxxxx; where h = any hydrophobic amino acid) that is highly conserved and a more precise representation of the FNIP repeat. At least eight Acanthamoeba polyphaga Mimivirus proteins and over 100 Dictyostelium proteins contain tandem arrays of the IP22 motif and its variants. cmbB also shares structural homology to YopM, from the plague bacterium Yersenia pestis.

Amino Acid Sequence↗

Protein linear indices of the 'macromolecular pseudograph alpha-carbon atom adjacency matrix' in bioinformatics. Part 1: prediction of protein stability effects of a complete set of alanine substitutions in Arc repressor.

A novel approach to bio-macromolecular design from a linear algebra point of view is introduced. A protein's total (whole protein) and local (one or more amino acid) linear indices are a new set of bio-macromolecular descriptors of relevance to protein QSAR/QSPR studies. These amino-acid level biochemical descriptors are based on the calculation of linear maps on Rn[f k(xmi):Rn-->Rn] in canonical basis. These bio-macromolecular indices are calculated from the kth power of the macromolecular pseudograph alpha-carbon atom adjacency matrix. Total linear indices are linear functional on Rn. That is, the kth total linear indices are linear maps from Rn to the scalar R[f k(xm):Rn-->R]. Thus, the kth total linear indices are calculated by summing the amino-acid linear indices of all amino acids in the protein molecule. A study of the protein stability effects for a complete set of alanine substitutions in the Arc repressor illustrates this approach. A quantitative model that discriminates near wild-type stability alanine mutants from the reduced-stability ones in a training series was obtained. This model permitted the correct classification of 97.56% (40/41) and 91.67% (11/12) of proteins in the training and test set, respectively. It shows a high Matthews correlation coefficient (MCC=0.952) for the training set and an MCC=0.837 for the external prediction set. Additionally, canonical regression analysis corroborated the statistical quality of the classification model (Rcanc=0.824). This analysis was also used to compute biological stability canonical scores for each Arc alanine mutant. On the other hand, the linear piecewise regression model compared favorably with respect to the linear regression one on predicting the melting temperature (tm) of the Arc alanine mutants. The linear model explains almost 81% of the variance of the experimental tm (R=0.90 and s=4.29) and the LOO press statistics evidenced its predictive ability (q2=0.72 and scv=4.79). Moreover, the TOMOCOMD-CAMPS method produced a linear piecewise regression (R=0.97) between protein backbone descriptors and tm values for alanine mutants of the Arc repressor. A break-point value of 51.87 degrees C characterized two mutant clusters and coincided perfectly with the experimental scale. For this reason, we can use the linear discriminant analysis and piecewise models in combination to classify and predict the stability of the mutant Arc homodimers. These models also permitted the interpretation of the driving forces of such folding process, indicating that topologic/topographic protein backbone interactions control the stability profile of wild-type Arc and its alanine mutants.

Alanine↗

Linear indices of the 'macromolecular graph's nucleotides adjacency matrix' as a promising approach for bioinformatics studies. Part 1: prediction of paromomycin's affinity constant with HIV-1 psi-RNA packaging region.

The design of novel anti-HIV compounds has now become a crucial area for scientists around the world. In this paper a new set of macromolecular descriptors (that are calculated from the macromolecular graph's nucleotide adjacency matrix) of relevance to nucleic acid QSAR/QSPR studies, nucleic acids' linear indices. A study of the interaction of the antibiotic Paromomycin with the packaging region of the HIV-1 psi-RNA has been performed as example of this approach. A multiple linear regression model predicted the local binding affinity constants [Log K (10(-4) M(-1))] between a specific nucleotide and the aforementioned antibiotic. The linear model explains more than 87% of the variance of the experimental Log K (R = 0.93 and s = 0.102 x 10(-4) M(-1)) and leave-one-out press statistics evidenced its predictive ability (q2 = 0.82 and s(cv) = 0.108 x 10(-4) M(-1)). The comparison with other approaches (macromolecular quadratic indices, Markovian Negentropies and 'stochastic' spectral moments) reveals a good behavior of our method.

Anti-Bacterial Agents↗

Affinity chromatography matures as bioinformatic and combinatorial tools develop.

Affinity chromatography has the reputation of a more expensive and less robust than other types of liquid chromatography. Furthermore, the technique is considered to stand a modest chance of large-scale purification of proteinaceous pharmaceuticals. This perception is changing because of the pressure for quality protein therapeutics, and the realization that higher returns can be expected when ensuring fewer purification steps and increased product recovery. These developments necessitated a rethinking of the protein purification processes and restored the interest for affinity chromatography. This liquid chromatography technique is designed to offer high specificity, being able to safely guide protein manufactures to successfully cope with the aforementioned challenges. Affinity ligands are distinguished into synthetic and biological. These can be generated by rational design or selected from ligand libraries. Synthetic ligands are generated by three methods. The rational method features the functional approach and the structural template approach. The combinatorial method relies on the selection of ligands from a library of synthetic ligands synthesized randomly. The combined method employs both methods, that is, the ligand is selected from an intentionally biased library based on a rationally designed ligand. Biological ligands are selected by employing high-throughput biological techniques, e.g. phage- and ribosome-display for peptide and microprotein ligands, in addition to SELEX for oligonucleotide ligands. Synthetic mimodyes and chimaeric dye-ligands are usually designed by rational approaches and comprise a chloro-triazinlyl scaffold. The latter substituted with various amino acids, carbocyclic, and heterocyclic groups, generates libraries from which synthetic ligands can be selected. A 'lead' compound may help to generating a 'focused' or 'biased' library. This can be designed by various approaches, e.g.: (i) using a natural ligand-protein complex as a template; (ii) applying the principle of complementarity to exposed residues of the protein structure; and (iii) mimicking directly a natural biological recognition interaction. Affinity ligands, based on the peptide structure, can be peptides, peptide-mimetic derivatives (<30 monomers) and microproteins (e.g. 25-200 monomers). Microprotein ligands are selected from biological libraries constructed of variegated protein domains, e.g. minibody, Kunitz, tendamist, cellulose-binding domain, scFv, Cytb562, zinc-finger, SpA-analogue (Z-domain).

Aldehyde Oxidoreductases↗

Venn analysis as part of a bioinformatic approach to prioritize expressed sequence tags from cardiac libraries.

OBJECTIVES: We needed to sort expressed sequence tags (ESTs) from human cardiac expression libraries. DESIGN AND METHODS: We annotated DNA sequence text files of 35,152 cardiac ESTs using our search and annotation tool called Multiblast.pl. We generated lists of the most prevalent ESTs in each library, and using a novel Venn tool, we grouped ESTs that were common to all or exclusive to particular libraries. RESULTS: Hypothetical protein KIAA0553 was expressed 120 times among 917 ESTs from an adult cardiac library (13.1%) compared only once among 8075 ESTs from fetal cardiac libraries (P < 10(-114)), this was confirmed using Northern analysis. We collated biochemical features of KIAA0553 and determined DNA polymorphism frequencies. We also used the Venn tool to specify genes that were uniquely expressed in hypertrophic cardiomyocytes. CONCLUSIONS: Annotating ESTs and sorting them using Venn analysis can help specify new candidate disease genes from the current lists of "hypothetical proteins".

Amino Acid Sequence↗

Bridging the gap between medical and bioinformatics: an ontological case study in colon carcinoma.

Ontological principles are needed in order to bridge the gap between medical and biological information in a robust and computable fashion. This is essential in order to draw inferences across the levels of granularity which span medicine and biology, an example of which include the understanding of the roles of tumor markers in the development and progress of carcinoma. Such information integration is also important for the integration of genomics information with the information contained in the electronic patient records in such a way that real time conclusions can be drawn. In this paper, we describe a large multi-granular datasource built by using ontological principles and focusing on the case of colon carcinoma.

Colonic Neoplasms↗

Bioinformatics and cellular signaling.

The understanding of cellular function requires an integrated analysis of context-specific, spatiotemporal data from diverse sources. Recent advances in describing the genomic and proteomic 'parts list' of the cell and deciphering the interrelationship of these parts are described, including genome-wide location analysis, standards for microarray data analysis, and two-hybrid and mass spectrometry approaches. This information is being collected and curated in databases such as the Alliance for Cellular Signaling (AfCS) Molecule Pages, which will serve as vital tools for the reconstruction and analysis of cellular signaling networks.

Cell Physiological Phenomena↗

Two novel presenilin 1 gene mutations connected with frontotemporal dementia-like clinical phenotype: genetic and bioinformatic assessment.

Mutations in the amyloid precursor protein (APP), presenilin 1 (PSEN1) and presenilin 2 (PSEN2) genes are associated with early-onset familial Alzheimer's disease (EOAD). There are several reports describing mutations in PSEN1 in cases with frontotemporal dementia (FTD). We identified two novel mutations in the PSEN1 gene: L226F and L424H. The first mutation was detected in a patient with a clinical diagnosis of FTD and a post-mortem diagnosis of AD. The second mutation is connected with a clinical phenotype of variant AD with strong FTD signs. In silico modeling revealed that the mutations, as well as mutations used for comparison (F177L and L424R), change the local structure, stability and/or properties of the transmembrane regions of the presenilin 1 protein (PS1). In contrast, a silent non-synonymous substitution F175S is eclipsed by external residues and has no influence on PS1 interfacial surface. We suggest that in silico analysis of PS1 substitutions can be used to characterize novel PSEN1 mutations, to discriminate between silent polymorphisms and a potential disease-causing mutation. We also propose that PSEN1 mutations should be considered in FTD patients with no MAPT mutations.

Adult↗

Bioinformatics, genomics and evolution of non-flagellar type-III secretion systems: a Darwinian perspective.

We review the biology of non-flagellar type-III secretion systems from a Darwinian perspective, highlighting the themes of evolution, conservation, variation and decay. The presence of these systems in environmental organisms such as Myxococcus, Desulfovibrio and Verrucomicrobium hints at roles beyond virulence. We review newly discovered sequence homologies (e.g., YopN/TyeA and SepL). We discuss synapomorphies that might be useful in formulating a taxonomy of type-III secretion. The problem of information overload is likely to be ameliorated by launch of a web site devoted to the comparative biology of type-III secretion ().

Amino Acid Sequence↗

Transcriptional and bioinformatic analysis of the 56.8 kb DNA region amplified in tandem repeats containing the penicillin gene cluster in Penicillium chrysogenum.

High penicillin-producing strains of Penicillium chrysogenum contain 6-14 copies of the three clustered structural biosynthetic genes, pcbAB, pcbC, and penDE [Barredo, J.L., Díez, B., Alvarez, E., Martín, J.F., 1989. Large amplification of a 35-kb DNA fragment carrying two penicillin biosynthetic genes in high penicillin producing strains of Penicillium chrysogenum. Curr. Genet. 16, 453-459; Smith, D.J., Bull, J.H., Edwards, J., Turner, G., 1989. Amplification of the isopenicillin N synthetase gene in a strain of Penicillium chrysogenum producing high levels of penicillin. Mol. Gen. Genet. 216, 492-497.] . The cluster is located in a 56.8 kb DNA region bounded by a conserved TGTAAA/T hexanucleotide that undergoes amplification in tandem repeats [Fierro, F., Barredo, J.L., Díez, B., Gutiérrez, S., Fernández, F.J., Martín, J.F., 1995. The penicillin gene cluster is amplified in tandem repeats linked by conserved hexanucleotide sequences. Proc. Natl. Acad. Sci. USA 92, 6200-6204; Newbert, R.W., Barton, B., Greaves, P., Harper, J., Turner, G., 1997. Analysis of a commercially improved Penicillium chrysogenum strain series: involvement of recombinogenic regions in amplification and deletion of the penicillin biosynthesis gene cluster. J. Ind. Microbiol. Biotechnol. 19, 18-27]. Transcriptional analysis of this amplified region (AR) revealed the presence of at least eight transcripts expressed in penicillin producing conditions. Three of them correspond to the known penicillin biosynthetic genes, pcbAB, pcbC, and penDE. To locate genes related to penicillin precursor formation, or penicillin transport and regulation we have sequenced and analyzed the 56.8 kb amplified region of P. chrysogenum AS-P-78, finding a total of 16 open reading frames. Two of these ORFs have orthologues of known function in the databases. Other ORFs showed similarities to specific domains occurring in different proteins and superfamilies which allowed to infer their probable function. ORF11 encodes a D-amino acid oxidase that might be responsible for the conversion of D-amino acids in the tripeptide L-alpha-aminoadipyl-L-cysteinyl-D-valine or other beta-lactam intermediates to deaminated by-products. ORF12 encodes a predicted protein with similarity to saccharopine dehydrogenases that seems to be related to biosynthesis of the penicillin precursor alpha-aminoadipic acid. A deletion mutant, P. chrysogenum npe10 lacking the entire AR including ORF12, shows a partial requirement of L-lysine for growth. ORF13 encodes a putative protein containing a Zn(II)2-Cys6 fungal-type DNA-binding domain, probably a transcriptional regulator. Although some of the ORFs in the AR may play roles in increasing penicillin production, none of the 13 ORFs other than pcbAB, pcbC, and penDE seem to be strictly indispensable for penicillin biosynthesis. The genes located in the P. chrysogenum AR have been compared with those found in the Aspergillus nidulans 50 kb DNA region adjacent to the penicillin gene cluster, showing no conservation between these two fungi.

Amino Acid Sequence↗