PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

The web server of IBM's Bioinformatics and Pattern Discovery group: 2004 update.

In this report, we provide an update on the services and content which are available on the web server of IBM's Bioinformatics and Pattern Discovery group. The server, which is operational around the clock, provides access to a large number of methods that have been developed and published by the group's members. There is an increasing number of problems that these tools can help tackle; these problems range from the discovery of patterns in streams of events and the computation of multiple sequence alignments, to the discovery of genes in nucleic acid sequences, the identification--directly from sequence--of structural deviations from alpha-helicity and the annotation of amino acid sequences for antimicrobial activity. Additionally, annotations for more than 130 archaeal, bacterial, eukaryotic and viral genomes are now available on-line and can be searched interactively. The tools and code bundles continue to be accessible from http://cbcsrv.watson.ibm.com/Tspd.html whereas the genomics annotations are available at http://cbcsrv.watson.ibm.com/Annotations/.

Anti-Infective Agents↗

The European Bioinformatics Institute's data resources: towards systems biology.

Genomic and post-genomic biological research has provided fine-grain insights into the molecular processes of life, but also threatens to drown biomedical researchers in data. Moreover, as new high-throughput technologies are developed, the types of data that are gathered en masse are diversifying. The need to collect, store and curate all this information in ways that allow its efficient retrieval and exploitation is greater than ever. The European Bioinformatics Institute's (EBI's) databases and tools have evolved to meet the changing needs of molecular biologists: since we last wrote about our services in the 2003 issue of Nucleic Acids Research, we have launched new databases covering protein-protein interactions (IntAct), pathways (Reactome) and small molecules (ChEBI). Our existing core databases have continued to evolve to meet the changing needs of biomedical researchers, and we have developed new data-access tools that help biologists to move intuitively through the different data types, thereby helping them to put the parts together to understand biology at the systems level. The EBI's data resources are all available on our website at http://www.ebi.ac.uk.

Computational Biology↗

E-MSD: an integrated data resource for bioinformatics.

The Macromolecular Structure Database (MSD) group (http://www.ebi.ac.uk/msd/) continues to enhance the quality and consistency of macromolecular structure data in the worldwide Protein Data Bank (wwPDB) and to work towards the integration of various bioinformatics data resources. One of the major obstacles to the improved integration of structural databases such as MSD and sequence databases like UniProt is the absence of up to date and well-maintained mapping between corresponding entries. We have worked closely with the UniProt group at the EBI to clean up the taxonomy and sequence cross-reference information in the MSD and UniProt databases. This information is vital for the reliable integration of the sequence family databases such as Pfam and Interpro with the structure-oriented databases of SCOP and CATH. This information has been made available to the eFamily group (http://www.efamily.org.uk/) and now forms the basis of the regular interchange of information between the member databases (MSD, UniProt, Pfam, Interpro, SCOP and CATH). This exchange of annotation information has enriched the structural information in the MSD database with annotation from wider sequence-oriented resources. This work was carried out under the 'Structure Integration with Function, Taxonomy and Sequences (SIFTS)' initiative (http://www.ebi.ac.uk/msd-srv/docs/sifts) in the MSD group.

Amino Acid Sequence↗

GraBCas: a bioinformatics tool for score-based prediction of Caspase- and Granzyme B-cleavage sites in protein sequences.

Caspases and granzyme B are proteases that share the primary specificity to cleave at the carboxyl terminal of aspartate residues in their substrates. Both, caspases and granzyme B are enzymes that are involved in fundamental cellular processes and play a central role in apoptotic cell death. Although various targets are described, many substrates still await identification and many cleavage sites of known substrates are not identified or experimentally verified. A more comprehensive knowledge of caspase and granzyme B substrates is essential to understand the biological roles of these enzymes in more detail. The relatively high variability in cleavage site recognition sequence often complicates the identification of cleavage sites. As of yet there is no software available that allows identification of caspase and/or granzyme with cleavage sites differing from the consensus sequence. Here, we present a bioinformatics tool 'GraBCas' that provides score-based prediction of potential cleavage sites for the caspases 1-9 and granzyme B including an estimation of the fragment size. We tested GraBCas on already known substrates and showed its usefulness for protein sequence analysis. GraBCas is available at http://wwwalt.med-rz.uniklinik-saarland.de/med_fak/humangenetik/software/index.html.

Caspases↗

RPBS: a web resource for structural bioinformatics.

RPBS (Ressource Parisienne en Bioinformatique Structurale) is a resource dedicated primarily to structural bioinformatics. It is the result of a joint effort by several teams to set up an interface that offers original and powerful methods in the field. As an illustration, we focus here on three such methods uniquely available at RPBS: AUTOMAT for sequence databank scanning, YAKUSA for structure databank scanning and WLOOP for homology loop modelling. The RPBS server can be accessed at http://bioserv.rpbs.jussieu.fr/ and the specific services at http://bioserv.rpbs.jussieu.fr/SpecificServices.html.

Computational Biology↗

SOAP-based services provided by the European Bioinformatics Institute.

SOAP (Simple Object Access Protocol) (http://www.w3.org/TR/soap) based Web Services technology (http://www.w3.org/ws) has gained much attention as an open standard enabling interoperability among applications across heterogeneous architectures and different networks. The European Bioinformatics Institute (EBI) is using this technology to provide robust data retrieval and data analysis mechanisms to the scientific community and to enhance utilization of the biological resources it already provides [N. Harte, V. Silventoinen, E. Quevillon, S. Robinson, K. Kallio, X. Fustero, P. Patel, P. Jokinen and R. Lopez (2004) Nucleic Acids Res., 32, 3-9]. These services are available free to all users from http://www.ebi.ac.uk/Tools/webservices.

Biotechnology↗

CAPweb: a bioinformatics CGH array Analysis Platform.

Assessing variations in DNA copy number is crucial for understanding constitutional or somatic diseases, particularly cancers. The recently developed array-CGH (comparative genomic hybridization) technology allows this to be investigated at the genomic level. We report the availability of a web tool for analysing array-CGH data. CAPweb (CGH array Analysis Platform on the Web) is intended as a user-friendly tool enabling biologists to completely analyse CGH arrays from the raw data to the visualization and biological interpretation. The user typically performs the following bioinformatics steps of a CGH array project within CAPweb: the secure upload of the results of CGH array image analysis and of the array annotation (genomic position of the probes); first level analysis of each array, including automatic normalization of the data (for correcting experimental biases), breakpoint detection and status assignment (gain, loss or normal); validation or deletion of the analysis based on a summary report and quality criteria; visualization and biological analysis of the genomic profiles and results through a user-friendly interface. CAPweb is accessible at http://bioinfo.curie.fr/CAPweb.

Chromosome Breakage↗

Identification of 17 Pseudomonas aeruginosa sRNAs and prediction of sRNA-encoding genes in 10 diverse pathogens using the bioinformatic tool sRNAPredict2.

sRNAs are small, non-coding RNA species that control numerous cellular processes. Although it is widely accepted that sRNAs are encoded by most if not all bacteria, genome-wide annotations for sRNA-encoding genes have been conducted in only a few of the nearly 300 bacterial species sequenced to date. To facilitate the efficient annotation of bacterial genomes for sRNA-encoding genes, we developed a program, sRNAPredict2, that identifies putative sRNAs by searching for co-localization of genetic features commonly associated with sRNA-encoding genes. Using sRNAPredict2, we conducted genome-wide annotations for putative sRNA-encoding genes in the intergenic regions of 11 diverse pathogens. In total, 2759 previously unannotated candidate sRNA loci were predicted. There was considerable range in the number of sRNAs predicted in the different pathogens analyzed, raising the possibility that there are species-specific differences in the reliance on sRNA-mediated regulation. Of 34 previously unannotated sRNAs predicted in the opportunistic pathogen Pseudomonas aeruginosa, 31 were experimentally tested and 17 were found to encode sRNA transcripts. Our findings suggest that numerous genes have been missed in the current annotations of bacterial genomes and that, by using improved bioinformatic approaches and tools, much remains to be discovered in 'intergenic' sequences.

Computational Biology↗

ApiDB: integrated resources for the apicomplexan bioinformatics resource center.

ApiDB (http://ApiDB.org) represents a unified entry point for the NIH-funded Apicomplexan Bioinformatics Resource Center (BRC) that integrates numerous database resources and multiple data types. The phylum Apicomplexa comprises numerous veterinary and medically important parasitic protozoa including human pathogenic species of the genera Cryptosporidium, Plasmodium and Toxoplasma. ApiDB serves not only as a database in its own right, but as a single web-based point of entry that unifies access to three major existing individual organism databases (PlasmoDB.org, ToxoDB.org and CryptoDB.org), and integrates these databases with data available from additional sources. Through the ApiDB site, users may pose queries and search all available apicomplexan data and tools, or they may visit individual component organism databases.

Animals↗

Reproducibility, bioinformatic analysis and power of the SAGE method to evaluate changes in transcriptome.

The serial analysis of gene expression (SAGE) method is used to study global gene expression in cells or tissues in various experimental conditions. However, its reproducibility has not yet been definitively assessed. In this study, we have evaluated the reproducibility of the SAGE method and identified the factors that affect it. The determination coefficient (R2 ) for the reproducibility of SAGE is 0.96. However, there are some factors that can affect the reproducibility of SAGE, such as the replication of concatemers and ditags, the number of sequenced tags and double PCR amplification of ditags. Thus, corrections for these factors must be made to ensure the reproducibility and accuracy of SAGE results. A bioinformatic analysis of SAGE data is also presented in order to eliminate these artifacts. Finally, the current study shows that increasing the number of sequenced tags improves the power of the method to detect transcripts and their regulation by experimental conditions.

Animals↗

In silico methods for evaluating human allergenicity to novel proteins: International Bioinformatics Workshop Meeting Report, 23-24 February 2005.

The ILSI Health and Environmental Sciences Institute (HESI) hosted an expert workshop 22-24 February 2005 in Mallorca, Spain, to review the state-of-the-science for conducting a sequence homology/bioinformatics evaluation in the context of a comprehensive allergenicity assessment for novel proteins, to obtain consensus on the value and role of bioinformatics in evaluating novel proteins, and to discuss the utility and methods of allergen-specific IgE testing in the diagnosis of food allergy. The workshop participants included over forty international experts from academia, industry, and government. The workshop was hosted by the HESI Protein Allergenicity Technical committee, which has established a long-term program whose mission is to advance the scientific understanding of the relevant parameters for characterizing the allergenic potential of novel proteins.

Allergens↗

Bioinformatics correctly identifies many type III secretion substrates in the plant pathogen Pseudomonas syringae and the biocontrol isolate P. fluorescens SBW25.

The plant pathogen Pseudomonas syringae causes disease by secreting a potentially large set of virulence proteins called effectors directly into host cells, their environment, or both, using a type III secretion system (T3SS). Most P. syringae effectors have a common upstream element called the hrp box, and their N-terminal regions have amino acids biases, features that permit their bioinformatic prediction. One of the most prominent biases is a positive serine bias. We previously used the truncated AvrRpt2(81-255) effector containing a serine-rich stretch from amino acids 81 to 100 as a T3SS reporter. Region 81 to 100 of this reporter does not contribute to the secretion or translocation of AvrRpt2 or to putative effector protein chimeras. Rather, the serine-rich region from the N-terminus of AvrRpt2 is important for protein accumulation in bacteria. Most of the N-terminal region (amino acids 15 to 100) is not essential for secretion in culture or delivery to plants. However, portions of this sequence may increase the efficiency of AvrRpt2 secretion, delivery to plants, or both. Two effectors previously identified with the AvrRpt2(81-255) reporter were secreted in culture independently of AvrRpt2, validating the use of the C terminus of AvrRpt2 as a T3SS reporter. Finally, using the reduced AvrRpt2(101-255) reporter, we confirmed seven predicted effectors from P. syringae pv. tomato DC3000, four from P. syringae pv. syringae B728a, and two from P. fluorescens SBW25.

Amino Acid Sequence↗

Bioinformatics-enabled identification of the HrpL regulon and type III secretion system effector proteins of Pseudomonas syringae pv. phaseolicola 1448A.

The ability of Pseudomonas syringae pv. phaseolicola to cause halo blight of bean is dependent on its ability to translocate effector proteins into host cells via the hypersensitive response and pathogenicity (Hrp) type III secretion system (T3SS). To identify genes encoding type III effectors and other potential virulence factors that are regulated by the HrpL alternative sigma factor, we used a hidden Markov model, weight matrix model, and type III targeting-associated patterns to search the genome of P. syringae pv. phaseolicola 1448A, which recently was sequenced to completion. We identified 44 high-probability putative Hrp promoters upstream of genes encoding the core T3SS machinery, 27 candidate effectors and related T3SS substrates, and 10 factors unrelated to the Hrp system. The expression of 13 of these candidate HrpL regulon genes was analyzed by real-time polymerase chain reaction, and all were found to be upregulated by HrpL. Six of the candidate type III effectors were assayed for T3SS-dependent translocation into plant cells using the Bordetella pertussis calmodulin-dependent adenylate cyclase (Cya) translocation reporter, and all were translocated. PSPPH1855 (ApbE-family protein) and PSPPH3759 (alcohol dehydrogenase) have no apparent T3SS-related function; however, they do have homologs in the model strain P. syringae pv. tomato DC3000 (PSPTO2105 and PSPTO0834, respectively) that are similarly upregulated by HrpL. Mutations were constructed in the DC3000 homologs and found to reduce bacterial growth in host Arabidopsis leaves. These results establish the utility of the bioinformatic or candidate gene approach to identifying effectors and other genes relevant to pathogenesis in P. syringae genomes.

Adenylyl Cyclases↗

Genomics of phytopathogenic fungi and the development of bioinformatic resources.

Genomic resources available to researchers studying phytopathogenic fungi are limited. Here, we briefly review the genomic and bioinformatic resources available and the current status of fungal genomics. We also describe a relational database containing sequences of expressed sequence tags (ESTs) from three phytopathogenic fungi, Blumeria graminis, Magnaporthe grisea, and Mycosphaerella graminicola, and the methods and underlying principles required for its construction. The database contains significant annotation for each EST sequence and is accessible at http://cogeme.ex.ac.uk. An easy-to-use interface allows the user to identify gene sequences by using simple text queries or homology searches. New querying functions and large sequence sets from a variety of phytopathogenic species will be incorporated in due course.

Cladosporium↗

Bringing the human genome and the revolution in bioinformatics to the medical school classroom: a case report from Washington University School of Medicine.

The human genome project is revolutionizing medical research and the practice of clinical medicine. To understand and participate in this revolution, physicians must be fluent in human genomics and bioinformatics. At Washington University School of Medicine (WUSM), the authors designed a module for teaching these skills to first-year students. The module uses clinical cases as a platform for accessing information stored in GenBank, Online Mendelian Inheritance in Man (OMIM), and PubMed databases at the National Center for Biotechnology Information (NCBI). This module, which is also designed to reinforce problem-solving skills, has been integrated into WUSM's first-year medical genetics course.

Computer-Assisted Instruction↗

Identification of NLRP3 and TIPE2 as asthma biomarkers via integrative bioinformatics and Mendelian randomization.

Asthma is a chronic inflammatory airway disease imposing a substantial global health burden. NLRP3 is an immune sensor involved in infection and cellular stress responses. Recent studies suggest that NLRP3 may be involved in the pathogenesis of asthma. We hypothesized that genetic variation in NLRP3 may contribute to asthma susceptibility. However, the causal relationship between NLRP3 and asthma still remains unclear. In this study, bioinformatics analysis using asthma data and R software was performed to identify NLRP3-related genes. We performed weighted gene co-expression network analysis to identify co-expressed genes, resulting in 12 candidate genes. Kyoto Encyclopedia of Genes and Genomes and Gene Ontology enrichment analyses were used to identify the functions of these candidate genes, revealing their involvement in cellular metabolism. Mendelian randomization analysis of the 12 candidate genes identified 2 biomarkers: NLRP3 and TNFAIP8L2 (TIPE2). We validated their diagnostic value for asthma using the GSE182503 dataset, with area under the curve values of 0.83 and 0.66 for NLRP3 and TIPE2, respectively. This project discusses how NLRP3 promotes asthma pathogenesis, whereas TIPE2 may alleviate it, and explores the potential interplay between them. NLRP3 and TIPE2 may serve as diagnostic biomarkers for asthma: NLRP3 may promote, whereas TIPE2 may alleviate asthma development. Both genes represent potential diagnostic biomarkers and therapeutic targets that warrant further functional investigation.

Asthma↗

Exploring the treatment of liver cancer with Gehua Hugan Gao based on bioinformatics, network pharmacology, and molecular docking.

Gehua Hugan Gao (GHHGG) is a traditional Chinese medicine paste that is chiefly used to treat liver cancer. However, the potential impact of GHHGG on liver cancer remains unclear. We explored how GHHGG treats liver cancer using bioinformatics, network pharmacology, and molecular docking. Network pharmacology included GHHGG active ingredients, predicted targets, predicted targets for liver cancer, and differential gene collection. A protein-protein interaction network was constructed using the Search Tool for the Retrieval of Interacting Genes/Proteins database, and crucial targets were ranked according to their degree values. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes analyses of liver cancer targets were followed by survival, differential analysis, and molecular docking. Venn diagrams show 123 predicted GHHGG targets for the treatment of hepatocellular carcinoma (HCC). Enrichment analysis showed that GHHGG treats HCC through multiple targets and pathways. We also found that estrogen receptor 1, cytochrome P450 3A4, cyclin-dependent kinase 4, type IIA topoisomerase, aurora kinase A, and cyclin E1 targets were closely associated with HCC development through survival and differential analyses. Molecular docking confirmed GHHGG's strong affinity for liver cancer targets. This study helps us understand GHHGG ingredients and targets for liver cancer treatment. To a certain extent, the molecular mechanism of GHHGG in the treatment of liver cancer has been elucidated, thus providing a theoretical basis.

Molecular Docking Simulation↗

Application of high-throughput computing in bioinformatics.

We describe the challenges faced when developing a Linux/PC-based cluster to apply bioinformatics algorithms to the rapidly increasing raw genomics data available. The calculations, which take around two months to complete, result in a powerful resource that can be used for data mining--most obviously for the human genome. Our current infrastructure consists of a 1314 node cluster with 1734 processors supporting both production and research. This paper highlights the problems in achieving high data throughput with such systems and shows that raw computer power is only one component of a complex problem.

Computational Biology↗