PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

[Necessity and usefulness of bioinformatic methods for microarray data analysis].

Data emerging from DNA microarray experiments are usually difficult to interpret. While the level of expression of several thousand genes can be measured in a single experiment, only a few dozen experiments are normally carried out, leading to data sets of very high dimensionality and low cardinality. The computational analysis of gene expression data makes significant usage of machine learning and statistical methods. Nevertheless, caution should be used in the blind adoption of these methods, as this usually leads to an over-interpretation of the expression profiles. The following presentation provides an overview of up-to-date principles of biostatistical analysis. A potential application for the analysis of high-dimensional expression profiles of prostate cancer is given.

Chromosome Aberrations↗

Panel of human cancer cell lines provides valuable database for drug discovery and bioinformatics.

Studies conducted at the US National Cancer Institute (NCI) and in our laboratory show that databases including the drug sensitivities of panels of many human cancer cell lines provide valuable information on the molecular pharmacology of anticancer drugs. We established a panel of 39 cell lines of various human cancers and developed a database of their chemosensitivities. Drugs were profiled in terms of their "fingerprints", patterns of differential activity against the cell lines. There was a significant correlation between a drug's fingerprint and its mode of action, as observed in the NCI panel of 60 cell lines. Therefore our cell-line panel is a powerful tool to predict the modes of action of new compounds. We have been using this system for drug discovery, coupled with various target-based drug screenings. We used the system to identify a novel DNA minor-groove binder, MS-247, which has inhibitory activity against topoisomerases I and II, and potent in vivo antitumor activity against various human cancer xenografts. We also discovered a potent novel telomerase inhibitor, FJ5002, by mining our database with the COMPARE algorithm, followed by experimental validation. We investigated the gene expression profiles of the cell lines by using DNA microarrays to find profiles determining cellular chemosensitivity and new targets for anticancer drugs. Our integrated database, including the chemosensitivities and gene expression profiles of the cell-line panel, could provide a basis for drug discovery and personalized therapy.

Databases, Factual↗

Bioinformatic analysis of primary endothelial cell gene array data illustrated by the analysis of transcriptome changes in endothelial cells exposed to VEGF-A and PlGF.

We recently published a review in this journal describing the design, hybridisation and basic data processing required to use gene arrays to investigate vascular biology (Evans et al. Angiogenesis 2003; 6: 93-104). Here, we build on this review by describing a set of powerful and robust methods for the analysis and interpretation of gene array data derived from primary vascular cell cultures. First, we describe the evaluation of transcriptome heterogeneity between primary cultures derived from different individuals, and estimation of the false discovery rate introduced by this heterogeneity and by experimental noise. Then, we discuss the appropriate use of Bayesian t-tests, clustering and independent component analysis to mine the data. We illustrate these principles by analysis of a previously unpublished set of gene array data in which human umbilical vein endothelial cells (HUVEC) cultured in either rich or low-serum media were exposed to vascular endothelial growth factor (VEGF)-A165 or placental growth factor (PlGF)-1(131). We have used Affymetrix U95A gene arrays to map the effects of these factors on the HUVEC transcriptome. These experiments followed a paired design and were biologically replicated three times. In addition, one experiment was repeated using serial analysis of gene expression (SAGE). In contrast to some previous studies, we found that VEGF-A and PlGF consistently regulated only small, non-overlapping and culture media-dependant sets of HUVEC transcripts, despite causing significant cell biological changes.

Cells, Cultured↗

Bioinformatic analysis of the TonB protein family.

TonB is a protein prevalent in a large number of Gram-negative bacteria that is believed to be responsible for the energy transduction component in the import of ferric iron complexes and vitamin B(12) across the outer membrane. We have analyzed all the TonB proteins that are currently contained in the Entrez database and have identified nine different clusters based on its conserved 90-residue C-terminal domain amino acid sequence. The vast majority of the proteins contained a single predicted cytoplasmic transmembrane domain; however, nine of the TonB proteins encompass a approximately 290 amino acid N-terminal extension homologous to the MecR1 protein, which is composed of three additional predicted transmembrane helices. The periplasmic linker region, which is located between the N-terminal domain and the C-terminal domain, is extremely variable both in length (22-283 amino acids) and in proline content, indicating that a Pro-rich domain is not a required feature for all TonB proteins. The secondary structure of the C-terminal domain is found to be well preserved across all families, with the most variable region being between the second alpha-helix and the third beta-strand of the antiparallel beta-sheet. The fourth beta-strand found in the solution structure of the Escherichia coli TonB C-terminal domain is not a well conserved feature in TonB proteins in most of the clusters. Interestingly, several of the TonB proteins contained two C-terminal domains in series. This analysis provides a framework for future structure-function studies of TonB, and it draws attention to the unusual features of several TonB proteins.

Amino Acid Sequence↗

A bioinformatics perspective on proteomics: data storage, analysis, and integration.

The field of proteomics is advancing rapidly as a result of powerful new technologies and proteomics experiments yield a vast and increasing amount of information. Data regarding protein occurrence, abundance, identity, sequence, structure, properties, and interactions need to be stored. Currently, a common standard has not yet been established and open access to results is needed for further development of robust analysis algorithms. Databases for proteomics will evolve from pure storage into knowledge resources, providing a repository for information (meta-data) which is mainly not stored in simple flat files. This review will shed light on recent steps towards the generation of a common standard in proteomics data storage and integration, but is not meant to be a comprehensive overview of all available databases and tools in the proteomics community.

Computational Biology↗

Molecular modeling and bioinformatical analysis of the antibacterial target enzyme MurA from a drug design perspective.

The enzyme MurA (UDP-N-acetylglucosamine enolpyruvyl transferase) catalyzes the first cytoplasmatic step in the synthesis of murein precursors. This function is of vital relevance for bacteria, and the enzyme therefore represents an important target protein for the development of novel antibacterial compounds. Several X-ray structures of liganded and un-liganded MurA have been published, which may be used for rational drug design. MurA, however, contains a highly flexible surface loop, which is involved in substrate and inhibitor binding. In the available X-ray structures, the conformation of this surface loop varies, depending on the presence or absence of ligands or substrate and probably also on the crystal packing. The uncertainty of the low-energy, or "resting state" conformation of this surface loop hampers the application of rational drug design to this class of enzymes. We have therefore performed an extensive molecular dynamics study of the enzyme in order to identify one or several low-energy conformers. The results indicate that, at least in some of the X-ray structures, the conformation of the flexible surface loop is influenced by crystallographic contacts. Furthermore, three partially helical foldamers of the surface loop are identified which may resemble the resting states of the enzyme or intermediate states that are "traversed" during the substrate binding process. Another, very important aspect for the development of novel antibacterial compounds is the inter- and intra-species variability of the target structure. We present a comparison of MurA sequences from 163 organisms which were analyzed under the aspects of enzyme mechanism, structure and drug design. The results allow us to identify the most promising binding sites for inhibitor interaction, which are present in MurA enzymes of most species and are expected to be insusceptible to resistance-inducing mutations.

Alkyl and Aryl Transferases↗

Mitochondrial ATP synthase: a bioinformatic approach reveals new insights about the roles of supernumerary subunits g and A6L.

The mitochondrial ATP synthase is a membrane protein complex which couples the proton gradient across the mitochondrial inner membrane to the synthesis of ATP from ADP+Pi. The complex is composed of essential subunits for its motor functions and supernumerary subunits, the roles of which remain to be elucidated. Subunits g and A6L are supernumerary subunits, and the specific roles of these subunits are still matters of debate. To gain insight into the functions of these two subunits, we carried out the alignment and the homolog search of the protein sequences of the subunits and found the following features: Subunit g appears to have isoforms in animals, and the transmembrane domain of the animal subunit g contains a completely conserved acidic residue in the middle of a helix on the conserved side of the transmembrane helix. This finding implicates the conserved acidic residue as important for the function of subunit g. The alignment of A6L protein sequences shows a conserved aromatic residue at the N-terminal domain with which the N-terminal MPQL sequence comprises a unique MPQLX4Ar motif that can signify the protein A6L. The conserved aromatic residue may also be important for the function of A6L.

Amino Acid Motifs↗

Soret spectral and bioinformatic approaches provide evidence for a critical role of the alpha -subunit in assembly of tetrameric hemoglobin.

Soret spectral contributions of the alpha-subunit heme pocket have been evaluated by performing static titrations of apohemoglobin A with CNProtohemin under varied experimental conditions. Increasing the temperature from 5 to 30 degrees C in 0.05 M potassium phosphate buffer, pH 7, resulted in a decreasingly prominent hypsochromic shifts reflecting altered the vinyl-globin interactions. Studies at 10 degrees C in over pH range of 6.7-8.0 revealed a profile for the spectral shifts approximating the side chain pK value (7.4) a histidyl residue. These overall spectral changes correspond to DeltaE of < or = 7 kJ/mol indicative of electrostatic noncovalent interactions. Further our current molecular modeling studies indicate that the spatial arrangement and critical noncovalent interactions of tyrosine 42 and histidine 45 (aromatic residues unique to the alpha-subunit) make significant contribution to the Soret spectra. Most interestingly, phylogenetic analyses have revealed the presence of a histidyl triad in the alpha-chain of all vertebrates that form heterotetramers.

Amino Acid Sequence↗

Comparative map and trait viewer (CMTV): an integrated bioinformatic tool to construct consensus maps and compare QTL and functional genomics data across genomes and experiments.

In the past few decades, a wealth of genomic data has been produced in a wide variety of species using a diverse array of functional and molecular marker approaches. In order to unlock the full potential of the information contained in these independent experiments, researchers need efficient and intuitive means to identify common genomic regions and genes involved in the expression of target phenotypic traits across diverse conditions. To address this need, we have developed a Comparative Map and Trait Viewer (CMTV) tool that can be used to construct dynamic aggregations of a variety of types of genomic datasets. By algorithmically determining correspondences between sets of objects on multiple genomic maps, the CMTV can display syntenic regions across taxa, combine maps from separate experiments into a consensus map, or project data from different maps into a common coordinate framework using dynamic coordinate translations between source and target maps. We present a case study that illustrates the utility of the tool for managing large and varied datasets by integrating data collected by CIMMYT in maize drought tolerance research with data from public sources. This example will focus on one of the visualization features for Quantitative Trait Locus (QTL) data, using likelihood ratio (LR) files produced by generic QTL analysis software and displaying the data in a unique visual manner across different combinations of traits, environments and crosses. Once a genomic region of interest has been identified, the CMTV can search and display additional QTLs meeting a particular threshold for that region, or other functional data such as sets of differentially expressed genes located in the region; it thus provides an easily used means for organizing and manipulating data sets that have been dynamically integrated under the focus of the researcher's specific hypothesis.

Adaptation, Physiological↗

Structure, circadian regulation and bioinformatic analysis of the unique sigma factor gene in Chlamydomonas reinhardtii.

In higher plants, the transcription of plastid genes is mediated by at least two types of RNA polymerase (RNAP); a plastid-encoded bacterial RNAP in which promoter specificity is conferred by nuclear-encoded sigma factors, and a nuclear-encoded phage-like RNAP. Green algae, however, appear to possess only the bacterial enzyme. Since transcription of much, if not most, of the chloroplast genome in Chlamydomonas reinhardtii is regulated by the circadian clock and the nucleus, we sought to identify sigma factor genes that might be responsible for this regulation. We describe a nuclear gene (RPOD) that is predicted to encode an 80 kDa protein that, in addition to a predicted chloroplast transit peptide at the N-terminus, has the conserved motifs (2.1- 4.2) diagnostic of bacterial sigma-70 factors. We also identified two motifs not previously recognized for sigma factors, adjacent PEST sequences and a leucine zipper, both suggested to be involved in protein-protein interactions. PEST sequences were also found in approximately 40% of sigma factors examined, indicating they may be of general significance. Southern blot hybridization and BLAST searches of the genome and EST databases suggest that RPODmay be the only sigma factor gene in C. reinhardtii. The levels of RPODmRNA increased 2- 3-fold in the mid-to-late dark period of light-dark cycling cells, just prior to, or coincident with, the peak in chloroplast transcription. Also, the dark-period peak in RPOD mRNA persisted in cells shifted to continuous light or continuous dark for at least one cycle, indicating that RPODis under circadian clock control. These results suggest that regulation of RPODexpression contributes to the circadian clock's control of chloroplast transcription.

Journal Article↗

Expression and prognosis of CXCL13 in uterine corpus endometrial carcinoma based on bioinformatics analysis.

OBJECTIVE: The biological significance of the chemokine ligand C-X-C motif chemokine ligand 13 (CXCL13) may play a significant role in the pathogenesis of uterine corpus endometrial carcinoma (UCEC). This study aims to identify and verify CXCL13 with predictive value for prognosis in UCEC. METHODS: CXCL13 mRNA expression differences were analyzed using R software in three independent datasets: one each from The Cancer Genome Atlas (TCGA) and two from the Gene Expression Omnibus (GEO), namely GSE17025 and GSE106191. The correlation between CXCL13 expression and prognosis was evaluated by Kaplan-Meier analysis. Univariate and multivariate Cox analyses were utilized to construct a prognostic nomogram. Tumor Immune Estimation Resource (TIMER) and the Tumor and Immune System Interaction Database (TISIDB) were employed to assess the relationship between CXCL13 and tumor immune infiltration. Coexpressed genes with CXCL13 were identified by the Spearman correlation analysis. A CXCL13 protein-protein interaction (PPI) network was constructed with the STRING website tool and hub genes were screened out. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genome (KEGG) analyses were performed with the "clusterProfiler" R package. Gene set enrichment analysis (GSEA) was used to identify underlying biological mechanisms. A drug-gene interaction network was constructed in the Comparative Toxicogenomics Database (CTD). RESULTS: High CXCL13 mRNA expression were validated in UCEC in the above three independent datasets. High CXCL13 expression was associated with favorable prognosis in UCEC. A nomogram for predicting the 1-, 3-, and 5-year survival probability in UCEC was construct based on CXCL13 expression and other clinical parameters. The use of Spearman correlation indicated certain correlation between CXCL13 and immune cells and immune checkpoint (ICP) genes. Seven hub genes were upregulated in UCEC, namely CXCL9, IFNG, CXCL10, CXCL11, GBP5, CCL18, and GZMB. The expression and prognostic relevance of CXCL9, IFNG, GBP5, and GZMB were in accordance with CXCL13. The main biological processes enriched were cytokine-cytokine receptor interaction and chemokine signaling pathway. CONCLUSIONS: The above comprehensive analyses suggest that CXCL13 may serve as a potential prognostic biomarker for UCEC, specifically for early-stage UCEC.

CXCL13↗

Bioinformatics.

Computer databases, networks and software tools are essential materials and methods for biomedical research and are involved in almost every aspect of disease gene mapping and positional cloning. Public databases of DNA and protein sequences and genetic and physical map information are increasing rapidly in size and complexity and are also improving in quality, comprehensiveness, interoperability and access. A new generation of software tools for navigating through the biomedical literature has become available. Programs for sequence homology searching and genetic map construction have become more sophisticated, yet easier to use. Global computer networks are bringing state-of-the-art capabilities to all.

Chromosome Mapping↗