PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Structural-functional bioinformatics: knowledge-based NMR interpretation.

This paper describes a knowledge-based approach to a problem of structural-functional bioinformatics, specifically the determination of protein structure through the automated analysis of NMR data. Highly successful results in carrying out sequence-specific assignments of residues from multidimensional NMR datasets has led us to automation of NOE dataset interpretation and a design for integrating these results with other protein structure and function analysis programs.

Computational Biology↗

Medical libraries, bioinformatics, and networked information: a coming convergence?

Libraries will be changed by technological and social developments that are fueled by information technology, bioinformatics, and networked information. Libraries in highly focused settings such as the health sciences are at a pivotal point in their development as the synthesis of historically diverse and independent information sources transforms health care institutions. Boundaries are breaking down between published literature and research data, between research databases and clinical patient data, and between consumer health information and professional literature. This paper focuses on the dynamics that are occurring with networked information sources and the roles that libraries will need to play in the world of medical informatics in the early twenty-first century.

Databases, Factual↗

Bacterial bioinformatics: pathogenesis and the genome.

As the number of completed microbial genome sequences continues to grow, there is a pressing need for the exploitation of this wealth of data through a synergistic interaction between the well-established science of bacteriology and the emergent discipline of bioinformatics. Antibiotic resistance and pathogenicity in virulent bacteria has become an increasing problem, with even the strongest drugs useless against some species, such as multi-drug resistant Enterococcus faecium and Mycobacterium tuberculosis. The global spread of Human Immunodeficiency Virus (HIV) and Acquired Immune Deficiency Syndrome (AIDS) has contributed to the re-emergence of tuberculosis and the threat from new and emergent diseases. To address these problems, bacterial pathogenicity requires redefinition as Koch's postulates become obsolete. This review discusses how the use of bacterial genomic information, and the in silico tools available at present, may aid in determining the definition of a current pathogen. The combination of both fields should provide a rapid and efficient way of assisting in the future development of antimicrobial therapies.

Acquired Immunodeficiency Syndrome↗

Large-scale open bioinformatics data resources.

The data explosion in bioinformatics is relentless. More and more genomes are being sequenced and many new types of datasets are being generated in large-scale projects. Integration and true open access to the data are still difficult issues, although they are gradually being addressed. Notably, certain fields have good standardization and interoperability, while others lag behind. This review summarizes the latest developments in genome and sequences databases, transcriptomics data (ESTs, ORESTES, full-length cDNAs), proteomics data (protein databases, protein structures, family and domain classification) as well as loosely integrated fields, such as microarray experiments, mutation databases and databases of regulatory regions and elements. The review attempts to resist simply summarizing what data are available, and aims to provide a critical look at some of the integration and access issues associated with several of these resources.

Computational Biology↗

Proteomics and bioinformatics approaches for identification of serum biomarkers to detect breast cancer.

BACKGROUND: Surface-enhanced laser desorption/ionization (SELDI) is an affinity-based mass spectrometric method in which proteins of interest are selectively adsorbed to a chemically modified surface on a biochip, whereas impurities are removed by washing with buffer. This technology allows sensitive and high-throughput protein profiling of complex biological specimens. METHODS: We screened for potential tumor biomarkers in 169 serum samples, including samples from a cancer group of 103 breast cancer patients at different clinical stages [stage 0 (n = 4), stage I (n = 38), stage II (n = 37), and stage III (n = 24)], from a control group of 41 healthy women, and from 25 patients with benign breast diseases. Diluted serum samples were applied to immobilized metal affinity capture Ciphergen ProteinChip Arrays previously activated with Ni2+. Proteins bound to the chelated metal were analyzed on a ProteinChip Reader Model PBS II. Complex protein profiles of different diagnostic groups were compared and analyzed using the ProPeak software package. RESULTS: A panel of three biomarkers was selected based on their collective contribution to the optimal separation between stage 0-I breast cancer patients and noncancer controls. The same separation was observed using independent test data from stage II-III breast cancer patients. Bootstrap cross-validation demonstrated that a sensitivity of 93% for all cancer patients and a specificity of 91% for all controls were achieved by a composite index derived by multivariate logistic regression using the three selected biomarkers. CONCLUSIONS: Proteomics approaches such as SELDI mass spectrometry, in conjunction with bioinformatics tools, could greatly facilitate the discovery of new and better biomarkers. The high sensitivity and specificity achieved by the combined use of the selected biomarkers show great potential for the early detection of breast cancer.

Adult↗

Bioinformatics and reanalysis of subtracted expressed sequence tags from the human ciliary body: Identification of novel biological functions.

PURPOSE: The ciliary body is largely known for its major roles in the regulation of aqueous humor secretion, intraocular pressure, and accommodation of the lens. In this review article we applied bioinformatics to re-examine hundreds of expressed sequence tags (ESTs) previously isolated by subtractive hybridization from a human ciliary body library [1]. The DNA sequences of these clones have been recently added to the web site of NEIBank. METHODS: DNA sequence comparisons of subtracted ESTs were performed against all entries in the last available release of the non-redundant database containing GenBank, EMBL, DDBJ and PDB sequences using the BlastN program accessed through NCBI's BLAST services on the internet (NCBI). Sequences were also compared and mapped using the Blast search program provided through the Internet by the Human Genome Project (UCSC). RESULTS: A total number of 284 independent ESTs were classified in 17 functional groups. Analysis of their relationships allowed to define the expression of five major groups of known genes: (i) protein synthesis, folding, secretion and degradation (20%); (ii) energy supply and biosynthesis (12%); (iii) contractility and cytoskeleton structure (6%); (iv) cellular signaling and cell cycle regulation (7%); and (v) nerve cell related tasks (2%), including neuropeptide processing and putative non-visual phototransduction and circadian rhythm control. The largest group contain unidentified sequences, a total of 105 sequences, accounting for 37% of ESTs. The unidentified sequences show similarity to genomic non-coding regions, or genes of unknown function. CONCLUSIONS: The most highly represented EST, correspond to myocilin, a gene involved in glaucoma. The data also confirms the secretory functions of the ciliary epithelium, and its high metabolism; the presence of a neuroendocrine peptidergic system presumably involved in the regulation of the intraocular pressure and/or aqueous humor secretion. Additional genes may be related to a non-visual phototransduction cascade and/or to circadian rhythms. Overall this initial group of subtracted ESTs can lead to uncover novel physiological functions of the ciliary body in normal and in disease, as well as novel candidate genes for ocular diseases.

Ciliary Body↗

[Disease-causing mutations versus neutral polymorphism: use of bioinformatics and DNA diagnosis].

Molecular genetic diagnostics is available for increasing number of genetically determined diseases. A wide spectrum of mutations can be detected by laboratory methods. A mutation can be defined as a change in a specific DNA sequence when compared with the reference sequence published in the gene database. However, in some cases it is difficult to distinguish if the detected sequence variant is a causal mutation or a neutral (polymorphic) variation without any effect on phenotype. The interpretation of rare sequence variants of unknown significance detected in disease-causing genes becomes an increasingly important problem. Further analysis on DNA and on protein levels with the use of bioinformatics are needed to reveal the effect of rare sequence variants. Inherited complex disorders, for example rare hereditary forms of cancer diseases, represent a challenge to molecular geneticists. The identification of exact causal mutation directly responsible for the development of the disease and for the assessment of disease risk resulting from this genetic variation has further implications. Predictive genetic diagnostics allows identify relatives at high risk of genetically determined disease and use of targeted preventive and therapeutic approaches. In severe cases it allows also prenatal or pre-implantation diagnostics.

DNA Mutational Analysis↗

A bioinformatics-based approach for the prediction and identification of novel proteins potentially involved in phosphorylation signalling pathways.

Together with the explosion in the availability of genome data of a number of organisms including human and mouse, various methods and programs for computational prediction of protein-coding genes and annotation of functional proteins have dramatically increased. For the last decade there has been intense interest in the role of protein phosphorylation which is involved in post-translation modification mechanisms critically regulating inter/intracellular communication, patho/physiological responses and homeostasis during many biological processes. In the present study a total of 202 functionally uncharacterized human full-coding cDNA sequences were investigated using a bioinformatics-based approach. Ten novel potential substrates for protein kinases have been identified which may play multiple roles in regulating intracellular phosphorylation signalling pathways. In addition, 5 of those may be involved in the human-only post-translation mechanism regulated by specific protein kinases. The data presented here therefore would greatly contribute toward the understanding of human molecular basis and cellular signalling networks.

Amino Acid Motifs↗

A novel bioinformatic strategy for unveiling hidden genome signatures of eukaryotes: self-organizing map of oligonucleotide frequency.

With the increasing amount of available genome sequences, novel tools are needed for comprehensive analysis of species-specific sequence characteristics for a wide variety of genomes. We used an unsupervised neural network algorithm, Kohonen's self-organizing map (SOM), to analyze di- and trinucleotide frequencies in 9 eukaryotic genomes of known sequences (a total of 1.2 Gb); S. cerevisiae, S. pombe, C. elegans, A. thaliana, D. melanogaster, Fugu, and rice, as well as P. falciparum chromosomes 2 and 3, and human chromosomes 14, 20, 21, and 22, that have been almost completely sequenced. Each genomic sequence with different window sizes was encoded as a 16- and 64-dimensional vector giving relative frequencies of di- and trinucleotides, respectively. From analysis of a total of 120,000 nonoverlapping 10-kb sequences and overlapping 100-kb sequences with a moving step size of 10 kb, derived from a total of the 1.2 Gb genomic sequences, clear species-specific separations of most sequences were obtained with the SOMs. The unsupervised algorithm could recognize, in most of the 120,000 10-kb sequences, the species-specific characteristics (key combinations of oligonucleotide frequencies) that are signature representations of each genome. Because the classification power is very high, the SOMs can provide fundamental bioinformatic strategies for extracting a wide range of genomic information that could not otherwise be obtained.

Animals↗

Bioinformatic approaches for identification and characterization of olfactomedin related genes with a potential role in pathogenesis of ocular disorders.

PURPOSE: To identify olfactomedin domain containing proteins, which are expressed in the eye and have similarity to myocilin, to test as potential candidates for eye diseases. Most of the mutations in myocilin causing primary open angle glaucoma are located in the olfactomedin domain. In vitro experiments demonstrated interaction between optimedin and myocilin through the conserved olfactomedin domains of the proteins in rats, and it was speculated that optimedin might have a role in the pathogenesis of ocular disorders. Hence, we aimed to identify myocilin related human proteins having conserved olfactomedin domains with potential to interact between them and examine the expression patterns in the eye by bioinformatics approaches. This endeavor would have the potential to identify new candidate genes for eye diseases in general and glaucoma in particular to be tested by wet-lab experiments. METHODS: Proteins with homology to myocilin were selected by BLASTp at the NCBI server. cDNA sequences and corresponding genomic contigs were retrieved. Pairwise BLAST was done to investigate the gene structure. The human EST database and NEIBank were searched against the selected cDNAs to look for tissue specific expression of the transcripts. RESULTS: The study led to the identification of three groups of proteins encoded by three different genes; Noelin 1 (9q34.3), Noelin 2 (19p13.2), and Noelin 3 (1p22) encompassing 45,575 bp, 82,679 bp, and 1,93,421 bp of the genomic sequence, respectively. Genomic structures, alternate usage of exons, and molecular evolution of the Noelins were determined. Similar structures of the genes, splicing patterns and high levels of homology shed light on the relatedness and molecular evolution of this group of olfactomedin related proteins. Strikingly, however, Noelin 1 and Noelin 3 were found to be expressed as multiple splice variants while only a single spliced transcript could be identified for Noelin 2. A human EST database search suggested the expression of all three Noelin genes in the brain but only two (Noelin 1 and Noelin 2) in the eye despite experimental evidence for expression of Noelin 3 in ocular tissue. Myocilin was determined to have similar levels (60-61%) of homology with all three Noelin gene products (Noelin 1_v1, Noelin 2_v1, and Noelin 3_v1) at the conserved olfactomedin domains. CONCLUSIONS: Mammalian Noelin 1 evolved from its precursor, followed by evolution of Noelin 3 and Noelin 2 by gene duplication events. Myocilin might have evolved from Noelin 2 by gene duplication followed by exon fusion. Noelin 1 and Noelin 2 could be tested as candidate genes for eye diseases based on their expressions in the eye and shared olfactomedin domains with Myocilin in the C-termini of the respective proteins.

Amino Acid Sequence↗

Support vector machine applications in bioinformatics.

The support vector machine (SVM) approach represents a data-driven method for solving classification tasks. It has been shown to produce lower prediction error compared to classifiers based on other methods like artificial neural networks, especially when large numbers of features are considered for sample description. In this review, the theory and main principles of the SVM approach are outlined, and successful applications in traditional areas of bioinformatics research are described. Current developments in techniques related to the SVM approach are reviewed which might become relevant for future functional genomics and chemogenomics projects. In a comparative study, we developed neural network and SVM models to identify small organic molecules that potentially modulate the function of G-protein coupled receptors. The SVM system was able to correctly classify approximately 90% of the compounds in a cross-validation study yielding a Matthews correlation coefficient of 0.78. This classifier can be used for fast filtering of compound libraries in virtual screening applications.

Algorithms↗

EYE on bioinformatics: dissecting complex disease traits in silico.

Bioinformatics has provided an unprecedented power and resource for us to decipher the enigma of complex diseases. It can reveal otherwise promiscuous information from the tremendous amount of data generated by the new, powerful and high-throughput technologies of genomics and proteomics. In this paper, we review the cutting edge developments in complex disease trait mapping, databases, computational gene recognition, gene function prediction, pathway reconstruction and disease classification by expression profiling, and computational modelling of living systems. Integration of all this knowledge and the different technologies, alongside cooperation between experts from different fields, will enhance our understanding of the molecular and mechanistic abnormalities in disease state, and greatly assist the rational development of effective therapies.

Chromosome Mapping↗

Distributed computing in bioinformatics.

This paper provides an overview of methods and current applications of distributed computing in bioinformatics. Distributed computing is a strategy of dividing a large workload among multiple computers to reduce processing time, or to make use of resources such as programs and databases that are not available on all computers. Participating computers may be connected either through a local high-speed network or through the Internet.

Computational Biology↗

Evolution of a Foundational Model of Physiology: symbolic representation for functional bioinformatics.

We describe the need for a Foundational Model of Physiology (FMP) as a reference ontology for "functional bioinformatics". The FMP is intended to support symbolic lookup, logical inference and mathematical analysis by integrating descriptive, qualitative and quantitative functional knowledge. The FMP will serve as a symbolic representation of biological functions initially pertaining to human physiology and ultimately extensible to other species. We describe the evolving architecture of the FMP, which is based on the ontological principles of the BioD biological description language and the Foundational Model of Anatomy (FMA).

Computational Biology↗

[Bioinformatics studies on photosynthetic system genes in cyanobacteria and chloroplasts].

This study compared homology of base sequences in genes encoding photosynthetic system proteins of cyanobacteria (Synechocystics sp. PCC6803, Nostoc sp. PCC7120) with these of chloroplasts (from Marchantia Polymorpha, Nicotiana tobacum, Oryza sativ, Euglena gracilis, Pinus thunbergii, Zea mays, Odentella sinesis, Cyanophora paradoxa, Porphyra purpurea and Arabidopsis thaliana) by BLAST method. While the gene sequence of Synechocystics sp. PCC6803 was considered as the criterion (100%) the homology of others were compared with it. Among the genes for photosystem I, psaC homology was the highest (90.14%) and the lowest was psaJ (52.24%). The highest ones were psbD (83.71%) for photosystem II, atpB (79.58%) for ATP synthase and petB (81.66%) for cytochrome b6/f complex. The lowest ones were psbN (49.70%) for photosystem II, atpF (26.69%) for ATP synthase and petA (55.27%) for cytochrome b6/f complex. Also, this paper discussed why the homology of gene sequences was the highest or the lowest. No report has been published and this bioinformatics research may provide some evidences for the origin and evolution of chloroplasts.

Chloroplast Proton-Translocating ATPases↗

Structural bioinformatics and QSAR analysis applied to the acetylcholinesterase and bispyridinium aldoximes.

The methods of bioinformatics, molecular modelling, and quantitative structure-activity relationships (QSARs) using regression and artificial neural network (ANN) analyses were applied to develop safer aldoxime antidotes against poisoning by organophosphorus (OP) agents with high, mean, and low aging rates. We start here from a molecular modelling of the mouse AChE at an atomistic level. Aim is to predict qualitatively the structural requirements of an aldoxime that shows an unique reactivating activity against the three classes of OPs. An antidotal action should occur by a three-site mechanism: the aldoxime groups of the first pyridinium ring should point towards the catalytic site, and the second pyridinium ring and its substituents should be anchored at the peripherical and anionic subsites. Based on this model, it is predicted that a suitable substituent is based on an arginine-like moiety. Then, an ANN-based QSAR analysis using a training set of aldoximes with known structure and activities was applied. Its input layer consisted of seven nodes: the group-membership descriptors that parameterize the type of the OP, the logarithms of the distribution coefficients at pH 7.4 and their squared term, the lowest unoccupied molecular orbital (LUMO) energies, the scaled molar refractions of the substituents, and their squared term. It was shown that the qualitative prediction made by molecular modelling can be quantified by an ANN prediction.

Acetylcholinesterase↗

Molecular characterisation of the SAND protein family: a study based on comparative genomics, structural bioinformatics and phylogeny.

The activities of vertebrate lysosomes are critical to many essential cellular processes. The yeast vacuole is analogous to the mammalian lysosome and is used as a tool to gain insights into vesicle mediated vacuolar/lysosome transport. The protein SAND, which does not contain a SAND domain (PFAM accession number PF01342), has recently been shown to function at the tethering/docking stage of vacuole fusion as a critical component of the vacuole SNARE complex. In this publication we have identified SAND in diverse eukaryotes, from single celled organisms such as the yeasts to complex multi-cellular chordates such as mammals. We have demonstrated subfamily divisions in the SAND proteins and show that in vertebrates, a duplication event gave rise to two SAND sequences. This duplication appears to have occurred during early vertebrate evolution and conceivably with the evolution of lysosomes. Using bioinformatics we predict a secondary structure, solvent accessibility profile and protein fold for the SAND proteins and determine conserved sequence motifs, present in all SAND proteins and those that are specific to subsets. A comprehensive evaluation of yeast and human functional studies in conjunction with our in silico analysis has identified potential roles for some of these motifs.

Amino Acid Sequence↗

Identification of protein pattern in kidney cancer using ProteinChip arrays and bioinformatics.

Tumor biology of renal cell carcinoma (RCC) is not very well understood, although many studies on molecular and cellular biology have been performed. It is accepted now that cancer research has to be performed also with proteomic tools, because proteins are the real actors in the genesis and progression of cancer. Therefore, we used a ProteinChip System(R) (SELDI) which is able to detect minute amounts of protein and moreover to analyze a complex protein pattern. We analyzed 37 cases of clear cell RCC as a training set including corresponding normal tissue. From all samples protein lysates were made and spotted directly on different chip surfaces (SAX2, WCX). After a washing procedure the arrays were analyzed in the ProteinChip Reader. All profiles were subjected to a bioinformatical analysis including normalization, clustering, rule extraction and rating. Defined rules (markers) were evaluated using a test set of 24 samples (13 tumor tissues and 11 normal kidney tissues). The generated rule base for the SAX2 surface showed a sensitivity of 100% and a specificity of 97.3%. For the WCX arrays the optimal rule base showed worse results. A combined rule base for SAX2 and WCX did not result in a higher sensitivity or specificity. Using the optimal rule base for the SAX2 chip in the test set, sensitivity and specificity reached 76.9% and 100%, respectively. The ProteinChip System represents a key technology for the rapid detection of cancer specific proteomic patterns. It is possible to identify clear cell renal cancer with high sensitivity and specificity from minimal amounts of cells.

Cell Line, Tumor↗