PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Structural biology and bioinformatics in drug design: opportunities and challenges for target identification and lead discovery.

Impressive progress in genome sequencing, protein expression and high-throughput crystallography and NMR has radically transformed the opportunities to use protein three-dimensional structures to accelerate drug discovery, but the quantity and complexity of the data have ensured a central place for informatics. Structural biology and bioinformatics have assisted in lead optimization and target identification where they have well established roles; they can now contribute to lead discovery, exploiting high-throughput methods of structure determination that provide powerful approaches to screening of fragment binding.

Computational Biology↗

The secreted protein discovery initiative (SPDI), a large-scale effort to identify novel human secreted and transmembrane proteins: a bioinformatics assessment.

A large-scale effort, termed the Secreted Protein Discovery Initiative (SPDI), was undertaken to identify novel secreted and transmembrane proteins. In the first of several approaches, a biological signal sequence trap in yeast cells was utilized to identify cDNA clones encoding putative secreted proteins. A second strategy utilized various algorithms that recognize features such as the hydrophobic properties of signal sequences to identify putative proteins encoded by expressed sequence tags (ESTs) from human cDNA libraries. A third approach surveyed ESTs for protein sequence similarity to a set of known receptors and their ligands with the BLAST algorithm. Finally, both signal-sequence prediction algorithms and BLAST were used to identify single exons of potential genes from within human genomic sequence. The isolation of full-length cDNA clones for each of these candidate genes resulted in the identification of >1000 novel proteins. A total of 256 of these cDNAs are still novel, including variants and novel genes, per the most recent GenBank release version. The success of this large-scale effort was assessed by a bioinformatics analysis of the proteins through predictions of protein domains, subcellular localizations, and possible functional roles. The SPDI collection should facilitate efforts to better understand intercellular communication, may lead to new understandings of human diseases, and provides potential opportunities for the development of therapeutics.

Cell Adhesion Molecules, Neuronal↗

Fast quantum search algorithms in protein sequence comparisons: quantum bioinformatics.

Quantum search algorithms are considered in the context of protein sequence comparison in bioinformatics. Given a sample protein sequence of length m (i.e., m residues), the problem considered is to find an optimal match in a large database containing N residues. Initially, Grover's quantum search algorithm is applied to a simple illustrative case-namely, where the database forms a complete set of states over the 2(m) basis states of a m qubit register, and thus is known to contain the exact sequence of interest. This example demonstrates explicitly the typical O(square root of [N]) speedup on the classical O(N) requirements. An algorithm is then presented for the (more realistic) case where the database may contain repeat sequences, and may not necessarily contain an exact match to the sample sequence. In terms of minimizing the Hamming distance between the sample sequence and the database subsequences the algorithm finds an optimal alignment, in O(square root of [N]) steps, by employing an extension of Grover's algorithm, due to Boyer et al. for the case when the number of matches is not a priori known.

Algorithms↗

Expression profiling and bioinformatic analyses of a novel stress-regulated multispanning transmembrane protein family from cereals and Arabidopsis.

Cold acclimation is a multigenic trait that allows hardy plants to develop efficient tolerance mechanisms needed for winter survival. To determine the genetic nature of these mechanisms, several cold-responsive genes of unknown function were identified from cold-acclimated wheat (Triticum aestivum). To identify the putative functions and structural features of these new genes, integrated genomic approaches of data mining, expression profiling, and bioinformatic predictions were used. The analyses revealed that one of these genes is a member of a small family that encodes two distinct groups of multispanning transmembrane proteins. The cold-regulated (COR)413-plasma membrane and COR413-thylakoid membrane groups are potentially targeted to the plasma membrane and thylakoid membrane, respectively. Further sequence analysis of the two groups from different plant species revealed the presence of a highly conserved phosphorylation site and a glycosylphosphatidylinositol-anchoring site at the C-terminal end. No homologous sequences were found in other organisms suggesting that this family is specific to the plant kingdom. Intraspecies and interspecies comparative gene expression profiling shows that the expression of this gene family is correlated with the development of freezing tolerance in cereals and Arabidopsis. In addition, several members of the family are regulated by water stress, light, and abscisic acid. Structure predictions and comparative genome analyses allow us to propose that the cor413 genes encode putative G-protein-coupled receptors.

Acclimatization↗

A complementary bioinformatics approach to identify potential plant cell wall glycosyltransferase-encoding genes.

Plant cell wall (CW) synthesizing enzymes can be divided into the glycan (i.e. cellulose and callose) synthases, which are multimembrane spanning proteins located at the plasma membrane, and the glycosyltransferases (GTs), which are Golgi localized single membrane spanning proteins, believed to participate in the synthesis of hemicellulose, pectin, mannans, and various glycoproteins. At the Carbohydrate-Active enZYmes (CAZy) database where e.g. glucoside hydrolases and GTs are classified into gene families primarily based on amino acid sequence similarities, 415 Arabidopsis GTs have been classified. Although much is known with regard to composition and fine structures of the plant CW, only a handful of CW biosynthetic GT genes-all classified in the CAZy system-have been characterized. In an effort to identify CW GTs that have not yet been classified in the CAZy database, a simple bioinformatics approach was adopted. First, the entire Arabidopsis proteome was run through the Transmembrane Hidden Markov Model 2.0 server and proteins containing one or, more rarely, two transmembrane domains within the N-terminal 150 amino acids were collected. Second, these sequences were submitted to the SUPERFAMILY prediction server, and sequences that were predicted to belong to the superfamilies NDP-sugartransferase, UDP-glycosyltransferase/glucogen-phosphorylase, carbohydrate-binding domain, Gal-binding domain, or Rossman fold were collected, yielding a total of 191 sequences. Fifty-two accessions already classified in CAZy were discarded. The resulting 139 sequences were then analyzed using the Three-Dimensional-Position-Specific Scoring Matrix and mGenTHREADER servers, and 27 sequences with similarity to either the GT-A or the GT-B fold were obtained. Proof of concept of the present approach has to some extent been provided by our recent demonstration that two members of this pool of 27 non-CAZy-classified putative GTs are xylosyltransferases involved in synthesis of pectin rhamnogalacturonan II (J. Egelund, B.L. Petersen, A. Faik, M.S. Motawia, C.E. Olsen, T. Ishii, H. Clausen, P. Ulvskov, and N. Geshi, unpublished data).

Amino Acid Motifs↗

Plant-based microarray data at the European Bioinformatics Institute. Introducing AtMIAMExpress, a submission tool for Arabidopsis gene expression data to ArrayExpress.

ArrayExpress is a public microarray repository founded on the Minimum Information About a Microarray Experiment (MIAME) principles that stores MIAME-compliant gene expression data. Plant-based data sets represent approximately one-quarter of the experiments in ArrayExpress. The majority are based on Arabidopsis (Arabidopsis thaliana); however, there are other data sets based on Triticum aestivum, Hordeum vulgare, and Populus subsp. AtMIAMExpress is an open-source Web-based software application for the submission of Arabidopsis-based microarray data to ArrayExpress. AtMIAMExpress exports data in MAGE-ML format for upload to any MAGE-ML-compliant application, such as J-Express and ArrayExpress. It was designed as a tool for users with minimal bioinformatics expertise, has comprehensive help and user support, and represents a simple solution to meeting the MIAME guidelines for the Arabidopsis community. Plant data are queryable both in ArrayExpress and in the Data Warehouse databases, which support queries based on gene-centric and sample-centric annotation. The AtMIAMExpress submission tool is available at http://www.ebi.ac.uk/at-miamexpress/. The software is open source and is available from http://sourceforge.net/projects/miamexpress/. For information, contact miamexpress@ebi.ac.uk.

Academies and Institutes↗

SPINE bioinformatics and data-management aspects of high-throughput structural biology.

SPINE (Structural Proteomics In Europe) was established in 2002 as an integrated research project to develop new methods and technologies for high-throughput structural biology. Development areas were broken down into workpackages and this article gives an overview of ongoing activity in the bioinformatics workpackage. Developments cover target selection, target registration, wet and dry laboratory data management and structure annotation as they pertain to high-throughput studies. Some individual projects and developments are discussed in detail, while those that are covered elsewhere in this issue are treated more briefly. In particular, this overview focuses on the infrastructure of the software that allows the experimentalist to move projects through different areas that are crucial to high-throughput studies, leading to the collation of large data sets which are managed and eventually archived and/or deposited.

Computational Biology↗

Grammatical inference in bioinformatics.

Bioinformatics is an active research area aimed at developing intelligent systems for analyses of molecular biology. Many methods based on formal language theory, statistical theory, and learning theory have been developed for modeling and analyzing biological sequences such as DNA, RNA, and proteins. Especially, grammatical inference methods are expected to find some grammatical structures hidden in biological sequences. In this article, we give an overview of a series of our grammatical approaches to biological sequence analyses and related researches and focus on learning stochastic grammars from biological sequences and predicting their functions based on learned stochastic grammars.

Algorithms↗

PFIT and PFRIT: bioinformatic algorithms for detecting glycosidase function from structure and sequence.

The identification of the enzymes involved in the metabolism of simple and complex carbohydrates presents one bioinformatic challenge in the post-genomic era. Here, we present the PFIT and PFRIT algorithms for identifying those proteins adopting the alpha/beta barrel fold that function as glycosidases. These algorithms are based on the observation that proteins adopting the alpha/beta barrel fold share positions in their tertiary structures having equivalent sets of atomic interactions. These are conserved tertiary interaction positions, which have been implicated in both structure and function. Glycosidases adopting the alpha/beta barrel fold share more conserved tertiary interactions than alpha/beta barrel proteins having other functions. The enrichment pattern of conserved tertiary interactions in the glycosidases is the information that PFIT and PFRIT use to predict whether any given alpha/beta barrel will function as a glycosidase or not. Using as a test set a database of 19 glycosidase and 45 nonglycosidase alpha/beta barrel proteins with low sequence similarity, PFIT and PFRIT can correctly predict glycosidase function for 84% of the proteins known to function as glycosidases. PFIT and PFRIT incorrectly predict glycosidase function for 25% of the nonglycosidases. The program PSI-BLAST can also correctly identify 84% of the 19 glycosidases, however, it incorrectly predicts glycosidase function for 50% of the nonglycosidases (twofold greater than PFIT and PFRIT). Overall, we demonstrate that the structure-based PFIT and PFRIT algorithms are both more selective and sensitive for predicting glycosidase function than the sequence-based PSI-BLAST algorithm.

Algorithms↗

Mutational and bioinformatics analysis of proline- and glycine-rich motifs in vesicular acetylcholine transporter.

The vesicular acetylcholine transporter (VAChT) contains six conserved sequence motifs that are rich in proline and glycine. Because these residues can have special roles in the conformation of polypeptide backbone, the motifs might have special roles in conformational changes during transport. Using published bioinformatics insights, the amino acid sequences of the 12 putative, helical, transmembrane segments of wild-type and mutant VAChTs were analyzed for propensity to form non-alpha-helical conformations and molecular notches. Many instances were found. In particular, high propensity for kinks and notches are robustly predicted for motifs D2, C and C'. Mutations in these motifs either increase or decrease Vmax for transport, but they rarely affect the equilibrium dissociation constants for ACh and the allosteric inhibitor, vesamicol. The near absence of equilibrium effects implies that the mutations do not alter the backbone conformation. In contrast, the Vmax effects demonstrate that the mutations alter the difficulty of a major conformational change in transport. Interestingly, mutation of an alanine to a glycine residue in motif C significantly increases the rates for reorientation across the membrane. These latter rates are deduced from the kinetics model of the transport cycle. This mutation is also predicted to produce a more flexible kink and tighter tandem notches than are present in wild-type. For the full set of mutations, faster reorientation rates correlate with greater predicted propensity for kinks and notches. The results of the study argue that conserved motifs mediate conformational changes in the VAChT backbone during transport.

Acetylcholine↗

Microarray and bioinformatic detection of novel and established genes expressed in experimental anti-Thy1 nephritis.

BACKGROUND: Microarray technology is a powerful tool that can probe the molecular pathogenesis of renal injury. In this present study microarray analysis was used to monitor serial changes in the renal transcriptome of a rat model of mesangial proliferative glomerulonephritis. Administration of anti-Thy1 antibody results in phases of acute mesangial injury (day 2), cell proliferation (day 5), matrix expansion (days 5 and 7), and subsequent healing (day 14). METHODS: Using Affymetrix (RAE230A) microarrays coupled with sequential primary biologic function-focused and secondary "baited" global cluster analysis, a cohort of established and putative novel modulators of mesangial cell turnover was identified. RESULTS: Cluster analysis of proliferative genes identified a number of gene expression profiles. The most striking pattern was increased gene expression at day 5, a cluster that included platelet-derived growth factor (PDGF), cyclins and transforming growth factor-beta (TGF-beta). The gene expression patterns identified by primary focused cluster analysis were used as bioinformatic bait and resulted in the identification of novel families of genes such as the S100 family. The expression of established and novel genes was confirmed using reverse transcription-polymerase chain reaction (RT-PCR). Next, in vivo gene expression was compared to PDGF-stimulated mesangial cells in vitro revealing similar patterns of dysregulation. CONCLUSION: Transcriptomic analysis defined both known and novel molecules involved in mesangial cell proliferation in vitro and in vivo and defined a panel of molecules that are potential contributors to mesangial cell dysfunction in glomerular disease.

Animals↗

Convergent analysis of cDNA and short oligomer microarrays, mouse null mutants and bioinformatics resources to study complex traits.

Gene expression data sets have recently been exploited to study genetic factors that modulate complex traits. However, it has been challenging to establish a direct link between variation in patterns of gene expression and variation in higher order traits such as neuropharmacological responses and patterns of behavior. Here we illustrate an approach that combines gene expression data with new bioinformatics resources to discover genes that potentially modulate behavior. We have exploited three complementary genetic models to obtain convergent evidence that differential expression of a subset of genes and molecular pathways influences ethanol-induced conditioned taste aversion (CTA). As a first step, cDNA microarrays were used to compare gene expression profiles of two null mutant mouse lines with difference in ethanol-induced aversion. Mice lacking a functional copy of G protein-gated potassium channel subunit 2 (Girk2) show a decrease in the aversive effects of ethanol, whereas preproenkephalin (Penk) null mutant mice show the opposite response. We hypothesize that these behavioral differences are generated in part by alterations in expression downstream of the null alleles. We then exploited the WebQTL databases to examine the genetic covariance between mRNA expression levels and measurements of ethanol-induced CTA in BXD recombinant inbred (RI) strains. Finally, we identified a subset of genes and functional groups associated with ethanol-induced CTA in both null mutant lines and BXD RI strains. Collectively, these approaches highlight the phosphatidylinositol signaling pathway and identify several genes including protein kinase C beta isoform and preproenkephalin in regulation of ethanol- induced conditioned taste aversion. Our results point to the increasing potential of the convergent approach and biological databases to investigate genetic mechanisms of complex traits.

Animals↗

Genome-wide screening of dioxin-responsive genes in fetal brain: bioinformatic and experimental approaches.

Many of the effects of dioxins, which are potent environmental pollutants and teratogens, are mediated through the aryl hydrocarbon receptor, also known as the dioxin receptor. The purpose of the present study was to characterize dioxin-responsive genes in a comprehensive manner using two complementary approaches: bioinformatic analysis and microarray analysis. First, we characterized the overall distribution of the cis-regulatory element for the dioxin-responsive element sequence (DRE) 'gcgtg' within putative promoter regions. We assembled the upstream sequences 10 kb from the transcription start site and evaluated their location and frequency in the human and mouse genomes. Second, we characterized the expression profile of mouse embryonic day 12 fetal brain exposed to 2,3,7,8-tetrarchlorodibenzo-p-dioxin. The distributions of 26,680 DREs among 2,843 human genes and 98,711 DREs among 18,541 mouse genes were examined. In both species, the DREs tended to be located close to the transcription start site. Forty genes exhibited significant induction or repression following dioxin exposure in fetal mice. The set of genes exhibited a strong functional coherence, with statistically significant enrichment in organogenesis and the DNA-dependent regulation of transcription, according to Gene Ontology annotations. In both humans and mice, DREs were preferentially distributed close to transcription start sites. Evolutionary conservation of this unique DRE distribution pattern suggests that DREs may be involved in transcriptional regulation. In mice, prenatal dioxin exposure altered the expression of 10 transcription factors, many of which have been documented to play a role in organogenesis. These genes may represent potential mediators of dioxin's effects in fetal tissues.

Animals↗

A new clan of CBM families based on bioinformatics of starch-binding domains from families CBM20 and CBM21.

Approximately 10% of amylolytic enzymes are able to bind and degrade raw starch. Usually a distinct domain, the starch-binding domain (SBD), is responsible for this property. These domains have been classified into families of carbohydrate-binding modules (CBM). At present, there are six SBD families: CBM20, CBM21, CBM25, CBM26, CBM34, and CBM41. This work is concentrated on CBM20 and CBM21. The CBM20 module was believed to be located almost exclusively at the C-terminal end of various amylases. The CBM21 module was known as the N-terminally positioned SBD of Rhizopus glucoamylase. Nowadays many nonamylolytic proteins have been recognized as possessing sequence segments that exhibit similarities with the experimentally observed CBM20 and CBM21. These facts have stimulated interest in carrying out a rigorous bioinformatics analysis of the two CBM families. The present analysis showed that the original idea of the CBM20 module being at the C-terminus and the CBM21 module at the N-terminus of a protein should be modified. Although the CBM20 functionally important tryptophans were found to be substituted in several cases, these aromatics and the regions around them belong to the best conserved parts of the CBM20 module. They were therefore used as templates for revealing the corresponding regions in the CBM21 family. Secondary structure prediction together with fold recognition indicated that the CBM21 module structure should be similar to that of CBM20. The evolutionary tree based on a common alignment of sequences of both modules showed that the CBM21 SBDs from alpha-amylases and glucoamylases are the closest relatives to the CBM20 counterparts, with the CBM20 modules from the glycoside hydrolase family GH13 amylopullulanases being possible candidates for the intermediate between the two CBM families.

Amino Acid Sequence↗

Interleukin-23 Receptor and Interleukin-17 Receptor A: Splice Variants, Isoforms and Their Relationship With Periodontitis-A Systematic Review and Bioinformatic Analysis.

This systematic review aimed to: (1) identify the splicing variants of IL23R and IL17RA reported in the literature; (2) perform a multiple alignment analysis to describe the isoforms of IL-23R and IL-17RA; and (3) compare the expression levels of IL-23R, IL-17RA, and their soluble isoforms (sIL-23R and sIL-17RA) in patients with periodontitis and periodontally healthy individuals. The study protocol followed PRISMA guidelines and was registered in PROSPERO (CRD420251267367). Six databases (PubMed, ScienceDirect, Scopus, Web of Science, EBSCO, and Google Scholar) were searched without restrictions on year or language. The descriptors used were: 'Interleukin-23 Receptor,' 'IL-23R,' 'Interleukin-17 Receptor A' 'IL-17RA,' 'Alternative Splicing,' 'Splice Variants,' 'Isoforms,' and 'Periodontitis.' The bioinformatics analysis was performed using CLUSTALW (V.1.83), InterPro and DeepTMHMM. Risk of bias was assessed with the QUIN and JBI tools for cross-sectional studies. Of 104 articles, four in vitro studies and eight cross-sectional studies were included. Qualitative analysis revealed that to date there are 32 splicing variants of the IL23R gene, while only one splicing variant has been reported for IL17RA. CLUSTALW, InterPro and DeepTMHMM analysis showed that these splicing variants result in 23 isoforms which can be soluble forms, complete intracellular peptides, truncated extracellular or intracellular peptides, or complete structures with truncated extracellular and/or intracellular domains. All studies had a low risk of bias. IL-23R and IL-17RA exhibit structural diversity resulting from alternative splicing, with IL-23R demonstrating significantly greater isoform complexity. However, the biological significance of these isoforms in periodontitis remains unclear and requires further investigation.

Humans↗

Structural bioinformatics-based design of selective, irreversible kinase inhibitors.

The active sites of 491 human protein kinase domains are highly conserved, which makes the design of selective inhibitors a formidable challenge. We used a structural bioinformatics approach to identify two selectivity filters, a threonine and a cysteine, at defined positions in the active site of p90 ribosomal protein S6 kinase (RSK). A fluoromethylketone inhibitor, designed to exploit both selectivity filters, potently and selectively inactivated RSK1 and RSK2 in mammalian cells. Kinases with only one selectivity filter were resistant to the inhibitor, yet they became sensitized after genetic introduction of the second selectivity filter. Thus, two amino acids that distinguish RSK from other protein kinases are sufficient to confer inhibitor sensitivity.

Amino Acid Sequence↗

Identification of secreted proteins of Mycobacterium tuberculosis by a bioinformatic approach.

Proteins secreted by Mycobacterium tuberculosis are usually targets of immune responses in the infected host. Here we describe a search for secreted proteins that combined the use of bioinformatics and phoA' fusion technology. The 3,924 proteins deduced from the M. tuberculosis genome were analyzed with several computer programs. We identified 52 proteins carrying an NH(2)-terminal secretory signal peptide but lacking additional membrane-anchoring moieties. Of these 52 proteins-the TM1 subgroup-only 7 had been previously reported to be secreted proteins. Our predictions were confirmed in 9 of 10 TM1 genes that were fused to Escherichia coli phoA', a marker of subcellular localization. These findings demonstrate that the systematic computer search described in this work identified secreted proteins of M. tuberculosis with high efficiency and 90% accuracy.

Algorithms↗

Comprehensive bioinformatic analysis of the specificity of human immunodeficiency virus type 1 protease.

Rapidly developing viral resistance to licensed human immunodeficiency virus type 1 (HIV-1) protease inhibitors is an increasing problem in the treatment of HIV-infected individuals and AIDS patients. A rational design of more effective protease inhibitors and discovery of potential biological substrates for the HIV-1 protease require accurate models for protease cleavage specificity. In this study, several popular bioinformatic machine learning methods, including support vector machines and artificial neural networks, were used to analyze the specificity of the HIV-1 protease. A new, extensive data set (746 peptides that have been experimentally tested for cleavage by the HIV-1 protease) was compiled, and the data were used to construct different classifiers that predicted whether the protease would cleave a given peptide substrate or not. The best predictor was a nonlinear predictor using two physicochemical parameters (hydrophobicity, or alternatively polarity, and size) for the amino acids, indicating that these properties are the key features recognized by the HIV-1 protease. The present in silico study provides new and important insights into the workings of the HIV-1 protease at the molecular level, supporting the recent hypothesis that the protease primarily recognizes a conformation rather than a specific amino acid sequence. Furthermore, we demonstrate that the presence of 1 to 2 lysine residues near the cleavage site of octameric peptide substrates seems to prevent cleavage efficiently, suggesting that this positively charged amino acid plays an important role in hindering the activity of the HIV-1 protease.

Algorithms↗