PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

ArrayQuest: a web resource for the analysis of DNA microarray data.

BACKGROUND: Numerous microarray analysis programs have been created through the efforts of Open Source software development projects. Providing browser-based interfaces that allow these programs to be executed over the Internet enhances the applicability and utility of these analytic software tools. RESULTS: Here we present ArrayQuest, a web-based DNA microarray analysis process controller. Key features of ArrayQuest are that (1) it is capable of executing numerous analysis programs such as those written in R, BioPerl and C++; (2) new analysis programs can be added to ArrayQuest Methods Library at the request of users or developers; (3) input DNA microarray data can be selected from public databases (i.e., the Medical University of South Carolina (MUSC) DNA Microarray Database or Gene Expression Omnibus (GEO)) or it can be uploaded to the ArrayQuest center-point web server into a password-protected area; and (4) analysis jobs are distributed across computers configured in a backend cluster. To demonstrate the utility of ArrayQuest we have populated the methods library with methods for analysis of Affymetrix DNA microarray data. CONCLUSION: ArrayQuest enables browser-based implementation of DNA microarray data analysis programs that can be executed on a Linux-based platform. Importantly, ArrayQuest is a platform that will facilitate the distribution and implementation of new analysis algorithms and is therefore of use to both developers of analysis applications as well as users. ArrayQuest is freely available for use at http://proteogenomics.musc.edu/arrayquest.html.

Algorithms↗

Kalign--an accurate and fast multiple sequence alignment algorithm.

BACKGROUND: The alignment of multiple protein sequences is a fundamental step in the analysis of biological data. It has traditionally been applied to analyzing protein families for conserved motifs, phylogeny, structural properties, and to improve sensitivity in homology searching. The availability of complete genome sequences has increased the demands on multiple sequence alignment (MSA) programs. Current MSA methods suffer from being either too inaccurate or too computationally expensive to be applied effectively in large-scale comparative genomics. RESULTS: We developed Kalign, a method employing the Wu-Manber string-matching algorithm, to improve both the accuracy and speed of multiple sequence alignment. We compared the speed and accuracy of Kalign to other popular methods using Balibase, Prefab, and a new large test set. Kalign was as accurate as the best other methods on small alignments, but significantly more accurate when aligning large and distantly related sets of sequences. In our comparisons, Kalign was about 10 times faster than ClustalW and, depending on the alignment size, up to 50 times faster than popular iterative methods. CONCLUSION: Kalign is a fast and robust alignment method. It is especially well suited for the increasingly important task of aligning large numbers of sequences.

Algorithms↗

Speeding disease gene discovery by sequence based candidate prioritization.

BACKGROUND: Regions of interest identified through genetic linkage studies regularly exceed 30 centimorgans in size and can contain hundreds of genes. Traditionally this number is reduced by matching functional annotation to knowledge of the disease or phenotype in question. However, here we show that disease genes share patterns of sequence-based features that can provide a good basis for automatic prioritization of candidates by machine learning. RESULTS: We examined a variety of sequence-based features and found that for many of them there are significant differences between the sets of genes known to be involved in human hereditary disease and those not known to be involved in disease. We have created an automatic classifier called PROSPECTR based on those features using the alternating decision tree algorithm which ranks genes in the order of likelihood of involvement in disease. On average, PROSPECTR enriches lists for disease genes two-fold 77% of the time, five-fold 37% of the time and twenty-fold 11% of the time. CONCLUSION: PROSPECTR is a simple and effective way to identify genes involved in Mendelian and oligogenic disorders. It performs markedly better than the single existing sequence-based classifier on novel data. PROSPECTR could save investigators looking at large regions of interest time and effort by prioritizing positional candidate genes for mutation detection and case-control association studies.

Algorithms↗

MAPPER: a search engine for the computational identification of putative transcription factor binding sites in multiple genomes.

BACKGROUND: Cis-regulatory modules are combinations of regulatory elements occurring in close proximity to each other that control the spatial and temporal expression of genes. The ability to identify them in a genome-wide manner depends on the availability of accurate models and of search methods able to detect putative regulatory elements with enhanced sensitivity and specificity. RESULTS: We describe the implementation of a search method for putative transcription factor binding sites (TFBSs) based on hidden Markov models built from alignments of known sites. We built 1,079 models of TFBSs using experimentally determined sequence alignments of sites provided by the TRANSFAC and JASPAR databases and used them to scan sequences of the human, mouse, fly, worm and yeast genomes. In several cases tested the method identified correctly experimentally characterized sites, with better specificity and sensitivity than other similar computational methods. Moreover, a large-scale comparison using synthetic data showed that in the majority of cases our method performed significantly better than a nucleotide weight matrix-based method. CONCLUSION: The search engine, available at http://mapper.chip.org, allows the identification, visualization and selection of putative TFBSs occurring in the promoter or other regions of a gene from the human, mouse, fly, worm and yeast genomes. In addition it allows the user to upload a sequence to query and to build a model by supplying a multiple sequence alignment of binding sites for a transcription factor of interest. Due to its extensive database of models, powerful search engine and flexible interface, MAPPER represents an effective resource for the large-scale computational analysis of transcriptional regulation.

Algorithms↗

Super paramagnetic clustering of protein sequences.

BACKGROUND: Detection of sequence homologues represents a challenging task that is important for the discovery of protein families and the reliable application of automatic annotation methods. The presence of domains in protein families of diverse function, inhomogeneity and different sizes of protein families create considerable difficulties for the application of published clustering methods. RESULTS: Our work analyses the Super Paramagnetic Clustering (SPC) and its extension, global SPC (gSPC) algorithm. These algorithms cluster input data based on a method that is analogous to the treatment of an inhomogeneous ferromagnet in physics. For the SwissProt and SCOP databases we show that the gSPC improves the specificity and sensitivity of clustering over the original SPC and Markov Cluster algorithm (TRIBE-MCL) up to 30%. The three algorithms provided similar results for the MIPS FunCat 1.3 annotation of four bacterial genomes, Bacillus subtilis, Helicobacter pylori, Listeria innocua and Listeria monocytogenes. However, the gSPC covered about 12% more sequences compared to the other methods. The SPC algorithm was programmed in house using C++ and it is available at http://mips.gsf.de/proj/spc. The FunCat annotation is available at http://mips.gsf.de. CONCLUSION: The gSPC calculated to a higher accuracy or covered a larger number of sequences than the TRIBE-MCL algorithm. Thus it is a useful approach for automatic detection of protein families and unsupervised annotation of full genomes.

Algorithms↗

HmtDB, a human mitochondrial genomic resource based on variability studies supporting population genetics and biomedical research.

BACKGROUND: Population genetics studies based on the analysis of mtDNA and mitochondrial disease studies have produced a huge quantity of sequence data and related information. These data are at present worldwide distributed in differently organised databases and web sites not well integrated among them. Moreover it is not generally possible for the user to submit and contemporarily analyse its own data comparing them with the content of a given database, both for population genetics and mitochondrial disease data. RESULTS: HmtDB is a well-integrated web-based human mitochondrial bioinformatic resource aimed at supporting population genetics and mitochondrial disease studies, thanks to a new approach based on site-specific nucleotide and aminoacid variability estimation. HmtDB consists of a database of Human Mitochondrial Genomes, annotated with population data, and a set of bioinformatic tools, able to produce site-specific variability data and to automatically characterize newly sequenced human mitochondrial genomes. A query system for the retrieval of genomes and a web submission tool for the annotation of new genomes have been designed and will soon be implemented. The first release contains 1255 fully annotated human mitochondrial genomes. Nucleotide site-specific variability data and multialigned genomes can be downloaded. Intra-human and inter-species aminoacid variability data estimated on the 13 coding for proteins genes of the 1255 human genomes and 60 mammalian species are also available. HmtDB is freely available, upon registration, at http://www.hmdb.uniba.it. CONCLUSION: The HmtDB project will contribute towards completing and/or refining haplogroup classification and revealing the real pathogenic potential of mitochondrial mutations, on the basis of variability estimation.

Computational Biology↗

MitoRes: a resource of nuclear-encoded mitochondrial genes and their products in Metazoa.

BACKGROUND: Mitochondria are sub-cellular organelles that have a central role in energy production and in other metabolic pathways of all eukaryotic respiring cells. In the last few years, with more and more genomes being sequenced, a huge amount of data has been generated providing an unprecedented opportunity to use the comparative analysis approach in studies of evolution and functional genomics with the aim of shedding light on molecular mechanisms regulating mitochondrial biogenesis and metabolism. In this context, the problem of the optimal extraction of representative datasets of genomic and proteomic data assumes a crucial importance. Specialised resources for nuclear-encoded mitochondria-related proteins already exist; however, no mitochondrial database is currently available with the same features of MitoRes, which is an update of the MitoNuc database extensively modified in its structure, data sources and graphical interface. It contains data on nuclear-encoded mitochondria-related products for any metazoan species for which this type of data is available and also provides comprehensive sequence datasets (gene, transcript and protein) as well as useful tools for their extraction and export. DESCRIPTION: MitoRes http://www2.ba.itb.cnr.it/MitoRes/ consolidates information from publicly external sources and automatically annotates them into a relational database. Additionally, it also clusters proteins on the basis of their sequence similarity and interconnects them with genomic data. The search engine and sequence management tools allow the query/retrieval of the database content and the extraction and export of sequences (gene, transcript, protein) and related sub-sequences (intron, exon, UTR, CDS, signal peptide and gene flanking regions) ready to be used for in silico analysis. CONCLUSION: The tool we describe here has been developed to support lab scientists and bioinformaticians alike in the characterization of molecular features and evolution of mitochondrial targeting sequences. The way it provides for the retrieval and extraction of sequences allows the user to overcome the obstacles encountered in the integrative use of different bioinformatic resources and the completeness of the sequence collection allows intra- and interspecies comparison at different biological levels (gene, transcript and protein).

Animals↗

Composition-based statistics and translated nucleotide searches: improving the TBLASTN module of BLAST.

BACKGROUND: TBLASTN is a mode of operation for BLAST that aligns protein sequences to a nucleotide database translated in all six frames. We present the first description of the modern implementation of TBLASTN, focusing on new techniques that were used to implement composition-based statistics for translated nucleotide searches. Composition-based statistics use the composition of the sequences being aligned to generate more accurate E-values, which allows for a more accurate distinction between true and false matches. Until recently, composition-based statistics were available only for protein-protein searches. They are now available as a command line option for recent versions of TBLASTN and as an option for TBLASTN on the NCBI BLAST web server. RESULTS: We evaluate the statistical and retrieval accuracy of the E-values reported by a baseline version of TBLASTN and by two variants that use different types of composition-based statistics. To test the statistical accuracy of TBLASTN, we ran 1000 searches using scrambled proteins from the mouse genome and a database of human chromosomes. To test retrieval accuracy, we modernize and adapt to translated searches a test set previously used to evaluate the retrieval accuracy of protein-protein searches. We show that composition-based statistics greatly improve the statistical accuracy of TBLASTN, at a small cost to the retrieval accuracy. CONCLUSION: TBLASTN is widely used, as it is common to wish to compare proteins to chromosomes or to libraries of mRNAs. Composition-based statistics improve the statistical accuracy, and therefore the reliability, of TBLASTN results. The algorithms used by TBLASTN are not widely known, and some of the most important are reported here. The data used to test TBLASTN are available for download and may be useful in other studies of translated search algorithms.

Algorithms↗

On the importance of being finished.

The publication of an increasing number of draft genome sequences presents problems that will only be resolved by improved search tools and by complete finishing of the sequences - and their deposition in publicly accessible databases.

Computational Biology↗

The GRID: the General Repository for Interaction Datasets.

We have developed a relational database, called the General Repository for Interaction Datasets (The GRID) to archive and display physical, genetic and functional interactions. The GRID displays data-rich interaction tables for any protein of interest, combines literature-derived and high-throughput interaction datasets, and is readily accessible via the web. Interactions parsed in The GRID can be viewed in graphical form with a versatile visualization tool called Osprey.

DNA, Fungal↗

Text-mining and information-retrieval services for molecular biology.

Text-mining in molecular biology -- defined as the automatic extraction of information about genes, proteins and their functional relationships from text documents -- has emerged as a hybrid discipline on the edges of the fields of information science, bioinformatics and computational linguistics. A range of text-mining applications have been developed recently that will improve access to knowledge for biologists and database annotators.

Computational Biology↗

Human proton/oligopeptide transporter (POT) genes: identification of putative human genes using bioinformatics.

The proton-dependent oligopeptide transporters (POT) gene family currently consists of approximately 70 cloned cDNAs derived from diverse organisms. In mammals, two genes encoding peptide transporters, PepT1 and PepT2 have been cloned in several species including humans, in addition to a rat histidine/peptide transporter (rPHT1). Because the Candida elegans genome contains five putative POT genes, we searched the available protein and nucleic acid databases for additional mammalian/human POT genes, using iterative BLAST runs and the human expressed sequence tags (EST) database. The apparent human orthologue of rPHT1 (expression largely confined to rat brain and retina) was represented by numerous ESTs originating from many tissues. Assembly of these ESTs resulted in a contiguous sequence covering approximately 95% of the suspected coding region. The contig sequences and analyses revealed the presence of several possible splice variants of hPHT1. A second closely related human EST-contig displayed high identity to a recently cloned mouse cDNA encoding cyclic adenosine monophosphate (cAMP)-inducible 1 protein (gi:4580995). This contig served to identify a PAC clone containing deduced exons and introns of the likely human orthologue (termed hPHT2). Northern analyses with EST clones indicated that hPHT1 is primarily expressed in skeletal muscle and spleen, whereas hPHT2 is found in spleen, placenta, lung, leukocytes, and heart. These results suggest considerable complexity of the human POT gene family, with relevance to the absorption and distribution of cephalosporins and other peptoid drugs.

Blotting, Northern↗

Precious cells contain precious information: strategies and pitfalls in expression analysis from a few cells.

Expression analysis, often encompassed in the term "functional genomics," is the link between physiology and molecular biology. Often, specific physiological changes in plant development are due to a limited number of genes, expressed exclusively in very few cells of an organ or organism. Compounding the situation, these physiological changes may also be transient. Therefore, searching for the responsible genes, though exciting and necessary to understand important processes, is hindered primarily by the scarcity of "precious" cells in the desired physiological state. Used judiciously, molecular methods such as reverse transcription polymerase chain reaction (RT-PCR), microarray analysis, or subtractive hybridization allow analysis of rare or special cells. Each of these methods has its advantages and pitfalls. Working with precious cells entails special biological strategies to avoid excessive work in obtaining the data and misinterpretation of it. To illustrate the logic and methods involved in working with precious cells-tissues, we describe how subtractive hybridization followed by expressed sequence tag (EST) sequencing can be used to search for a few genes specific to a few available cells.

Base Sequence↗

Generating unigene collections of expressed sequence tag sequences for use in mass spectrometry identification.

Expressed sequence tag sequences remain the largest resource of DNA sequence for most organisms despite recent advances in genome sequencing. These sequences are short, fragmented versions of the expressed genes. By DNA sequence assembly, the fragments can be assembled into contiguous DNA sequences that are better suited for protein identification by mass spectrometry.

Cluster Analysis↗

The biology of the post-genomic era: the proteomics.

The complete identification of coding sequences in a number of species has led to announce the beginning of the post-genomic era, new tools have become available to study complex phenomena in biological systems. Rapid advances in genomic sequencing and bioinformatics have established the field of genomics to investigate thousands genes' activity through mRNA display. However, recent studies have demonstrated a lack of correlation between the transcriptional profiles and the actual protein levels in cells, so investigation of the expressed part of the genome is also required to link genomic data to biological function. It is possible that evolutional development occured by increasing complexity of regulation processes at the level of RNA and protein molecules instead of simple increase in gene number, so investigation of proteins and protein complexes became important fields of our post-genomic era. High-resolution two-dimensional gels combined with sensitive mass spectrometry can reveal virtually all proteins present in cells opening new insights into functions of cells, tissues and whole organisms.

Animals↗

Molecular cloning of two novel transmembrane ligands for Eph-related kinases (LERKS) that are related to LERK-2.

A search of the nucleic acid database of expressed sequence tags (ESTs) revealed several partial cDNA sequences that could encode proteins homologous to the ligands for Eph-related kinases (LERKs). Oligonucleotides designed from the ESTs were used to probe a human brain cDNA library and obtain overlapping clones that encoded two different novel LERKS (NLERK-1 and NLERK-2). NLERK-1 and NLERK-2 are most closely related to human LERK-2/Elk-ligand and they form a subclass of LERKs that contain a transmembrane domain and a conserved cytoplasmic domain. Full-length NLERK-1 was expressed as a glycosylated membrane protein in COS cells and was not secreted into the medium. Full-length NLERK-2 was similarly expressed in COS cells but both membrane-bound and a truncated, proteolytically-released form were detected. Engineered forms of both NLERK-1 and NLERK-2 lacking transmembrane and cytoplasmic domains were also expressed in COS cells and each was detected in the extracellular medium.

Amino Acid Sequence↗

An integrated strategy for the optimization of microarray data interpretation.

The completion of a microarray experiment represents just a starting point toward understanding the biology of interest. A follow-up strategy is needed to fully elucidate the functional significance of microarray-derived measurements of differential expression. Given the fact that no single approach can fully unravel the fundamental biology that is typically quite complex, the follow-up strategy must be integrated at multiple levels encompassing bioinformatics, genomics, and proteomics. In this review, we discuss an integrative approach, which can be used to prioritize microarray-derived candidate genes, define their functions, and place them in the context of the biological system being studied.

Animals↗